BACK Mascot image.

MA0 0

Math Prepa for University

Every lesson so far, in one document · 5 chapters

back to the contents

Lesson 1

Numbers, and Sets

Taught

We note early that the notes below and onwards assume the basic algebra and arithmetic of a GCSE math course; everything past that point is built up here from the start.

The Number Line

Most of what follows rests on two ideas: the number line, which is where numbers live, and the set, which is how we talk about a collection of them at once.

We start with the number line. Here is an infinite line.

Figure 1.1. An infinite line.

The arrowheads are there to say that the line is meant to extend in both directions, to infinity. Now let us fix a point on this line (can be any point, but we do normally like it being in the middle), and label it 00.

0
Figure 1.2. A zero point fixed on the line.

This point represents the number zero, and it is the point against which every other number is measured. Numbers sitting to the left of 00 are called negative, numbers sitting to the right are called positive, and 00 itself is neither.

One thing is still missing; we can now say which side of 00 a point lies on, but not how far along it lies, so let us finally fix a unit of length.

01one unit
Figure 1.3. A unit of length fixed on the line.

This unit of length is what we use, among other things, to compare how far the other numbers sit from zero.

Definition 1.1 (The Number Line).

The infinite line above, together with its fixed zero point and its fixed unit of length, is the (real) number line. The numbers represented by the points of the number line are called real numbers, and the collection of all of them is written R\RR.

A few real numbers, and where they sit:

−π−2.2log(1/30)01√2e
Figure 1.4. Seven real numbers on the number line.

The Natural Numbers

The natural numbers are the numbers we count with. Counting is done by incrementing a quantity, one object at a time, and on the number line an increment is a move to the right by the unit length. The distance we have moved from zero is then the number of times we moved, which is exactly the number of objects counted.

Definition 1.2 (Natural Numbers).

The natural numbers (sometimes called whole numbers) are the numbers represented by those points of the number line that can be reached by starting at 00 and moving right by a whole unit length some number of times.

012345…
Figure 1.5. The natural numbers. The point 00 is drawn hollow because we do not count it as one of them.

Remark (A Convention).

Whether 00 counts as a natural number is a matter of convention; but my personal preference is we do not count it, so N={1,2,3,…},\NN = \{1, 2, 3, \ldots\}, and when we do want zero alongside them we write N0={0,1,2,…}.\NN_0 = \{0, 1, 2, \ldots\}.

The Integers

The natural numbers count, but they cannot measure the difference between two counts; if an increment in quantity is a move to the right by the unit length, then a decrement is a move to the left by the unit length, and allowing both directions gives the integers.

Definition 1.3 (Integers).

The integers are the numbers represented by those points of the number line that can be reached by starting at 00 and moving in either direction by a whole unit length some number of times. The collection of all of them is written Z\ZZ.

−5−4−3−2−1012345
Figure 1.6. The integers.

Closure

One property that ordinary arithmetic assumes without comment has not been mentioned yet.

Definition 1.4 (Closure).

A collection of numbers is closed under an operation if carrying out that operation on any two of its members again produces a member of that same collection.

The real numbers are closed under addition, subtraction, multiplication and division: the sum, difference, product and ratio of any two real numbers is again a real number. Division is the one case needing a footnote, and we come back to it below.

Smaller collections need not behave so well: If aa and bb are natural numbers then so are a+ba + b and abab, and if aa and bb are integers then so are a+ba + b, a−ba - b and abab; but by our definition the naturals are not closed under subtraction, since 1−21 - 2 is not a natural number, and the integers are not closed under division, since 1/21/2 is not an integer. Each failure is a reason to enlarge the collection we are working with.

Problem 1.1.

Which of the four operations is the collection of negative integers closed under?

The Rational Numbers

Subtraction led us to the integers; division leads to the next enlargement. Where an integer is what we get by laying whole units end to end, a rational number is what we get by first cutting the unit into equal segments.

Definition 1.5 (Rational Number).

A rational number is a real number that can be written as a fraction whose numerator is an integer and whose denominator is a nonzero integer; that is, xx is rational if and only if there are integers aa and bb with b≠0b \neq 0 such that x=a/b.x = a/b. The collection of all rational numbers is written Q\QQ.

Placing them on the line takes a little more work than placing the integers did. Given a positive integer bb, the number 1/b1/b is the length obtained by breaking the unit length into bb equal parts. For b=3b = 3:

11/31/31/3
Figure 1.7. The unit length broken into three equal parts.

The numbers of the form a/ba/b with aa an integer are then those reached by starting at 00 and moving left or right by 1/b1/b units some number of times, with aa deciding both the direction and the number of steps.

−2−1012−8/3−7/3−6/3−5/3−4/3−3/3−2/3−1/30/31/32/33/34/35/36/37/38/3
Figure 1.8. Thirds on the number line. Every integer is a whole number of thirds, three for each unit.

Nothing new happens when bb is negative, because then −b-b is positive and a/b=−a/−b,a/b = -a/-b, so the picture for a negative denominator is the picture for a positive one.

Remark (Why the denominator cannot be zero).

In terms of the picture above, a/0a/0 asks us to break the unit into zero equal parts, which does nothing at all: there is no piece left to step by, so no point on the line gets named. The same objection can be made in one line. Dividing by bb is meant to undo multiplying by bb, so if a/0a/0 were a number qq we would need a=q⋅0=0a = q \cdot 0 = 0. If a≠0a \neq 0 there is then no such qq at all, and if a=0a = 0 then every qq qualifies. Either there is no answer or there is no way to choose between the answers, so we leave a/0a/0 undefined.

The Rationals are Closed

The rationals are closed under addition, subtraction and multiplication as the integers are, and unlike the integers they are also closed under division. This is quick to see.

Let xx and yy be rational. By the definition we may write x=a/bx = a/b and y=c/dy = c/d, where aa, bb, cc, dd are integers and b≠0b \neq 0, d≠0d \neq 0. Then

x+y=ab+cd=ad+bcbd.x + y = \frac{a}{b} + \frac{c}{d} = \frac{ad + bc}{bd}.

The integers are closed under multiplication, so adad, bcbc and bdbd are integers; they are closed under addition, so ad+bcad + bc is an integer; and bd≠0bd \neq 0, since neither bb nor dd is zero. So x+yx + y is an integer over a nonzero integer, which is to say it is rational. The same follows for the rest. For the difference and the product,

x−y=ad−bcbd,xy=acbd,x - y = \frac{ad - bc}{bd}, \qquad xy = \frac{ac}{bd},

and both are again an integer over a nonzero integer. For the quotient, suppose also that y≠0y \neq 0, so that c≠0c \neq 0; then

xy=ab⋅dc=adbc,\frac{x}{y} = \frac{a}{b} \cdot \frac{d}{c} = \frac{ad}{bc},

and bc≠0bc \neq 0, so x/yx/y is rational as well.

Problem 1.2.

Where exactly in the argument above was it used that bb and dd are nonzero? Would anything go wrong if aa or cc were zero?

Sets

The number line says what a single number is; to speak about many at once (the naturals, the integers, the solutions of an equation) we need the second idea.

Definition 1.6 (Set).

A set is a collection of distinguishable objects, mathematical or otherwise.

The objects need not be numbers. The set of all animals in this building called Bob is a perfectly good set, and so is the set of letters appearing in the word “banana”, which is the collection consisting of bb, aa and nn.

Definition 1.7 (Membership).

The objects collected in a set are its elements. If aa is one of the objects collected in the set AA, we say aa is an element of AA, or that aa belongs to AA, and write a∈A.a \in A. If aa is not one of them we write a∉Aa \notin A, read ”aa is not in AA”.

Example 1.8.

Let BB be the set of animals in this building called Bob. If the cat is called Bob then the cat∈B\text{the cat} \in B, and if the dog is called Rex then the dog∉B\text{the dog} \notin B. Numbers behave the same way: −2∈Z-2 \in \ZZ, while −2∉N-2 \notin \NN.

Subsets

In the example above, every animal called Bob is also an animal in the building. That situation is common enough to deserve a name: given two sets AA and BB, it may happen that every element of AA is also an element of BB, which is to say that whenever x∈Ax \in A we also have x∈Bx \in B.

Definition 1.9 (Subset).

A set AA is included in a set BB, or is a subset of BB, if every element of AA is also an element of BB. We write A⊂B.A \subset B.

If every element of AA lies in BB and every element of BB lies in AA, then the two collections have exactly the same members, and we call them equal: A=BA = B precisely when A⊂BA \subset B and B⊂AB \subset A.

Remark.

Sometimes we want to say that AA is contained in BB and that at least one element of BB is missing from AA, so that the two are not equal. Then we call AA a proper subset of BB and write A⊊BA \subsetneq B.

The collections of numbers built up in the first half of this chapter now line up in a chain,

N⊊N0⊊Z⊊Q⊊R,\NN \subsetneq \NN_0 \subsetneq \ZZ \subsetneq \QQ \subsetneq \RR,

and every inclusion in it is proper: 00 lies in N0\NN_0 but not in N\NN, −1-1 lies in Z\ZZ but not in N0\NN_0, 1/21/2 lies in Q\QQ but not in Z\ZZ, and 2\sqrt{2} lies in R\RR but not in Q\QQ; we do not prove the last of these here, but Exercise 1.6 does.

Remark (Decorating the symbols).

Three decorations on these symbols come up constantly. A star removes zero, so Z∗\ZZ^{*}, Q∗\QQ^{*} and R∗\RR^{*} are the integers, rationals and reals apart from 00. A subscript plus keeps the non-negative ones and a subscript minus the non-positive ones, giving Z+\ZZ_{+}, Q+\QQ_{+}, R+\RR_{+} and Z−\ZZ_{-}, Q−\QQ_{-}, R−\RR_{-}. The two can be combined: R+∗\RR_{+}^{*} is the set of strictly positive real numbers.

It is worth pointing out how much the symbol ⊂\subset resembles the symbol <<, and where the resemblance stops. Given two numbers aa and bb, exactly one of a<ba < b, a=ba = b, b<ab < a holds; two numbers cannot be different and yet fail to be comparable. Sets are not like this. Two sets can differ without either one being contained in the other.

Example 1.10.

Let A={1,2}A = \{1, 2\} and B={2,3}B = \{2, 3\}. Then A≠BA \neq B, but A⊄BA \not\subset B, since 1∈A1 \in A and 1∉B1 \notin B, and also B⊄AB \not\subset A, since 3∈B3 \in B and 3∉A3 \notin A. Neither set comes first.

Remark.

As in the example, a slash through the symbol denies it: A⊄BA \not\subset B says that AA is not a subset of BB, which is to say that at least one element of AA escapes BB. It is not used often.

Problem 1.3.

Is N∈N0\NN \in \NN_0? Is N⊂N0\NN \subset \NN_0? These are not the same question.

Problem 1.4.

Is

3+43−4∈Q ?\frac{\sqrt{3} + \sqrt{4}}{\sqrt{3} - \sqrt{4}} \in \QQ\,?

Listing and Describing

Describing a set in a sentence, as we have been doing, is exact but slow. Two shorter notations do most of the work. The first is simply to list the elements, separated by commas and enclosed in curly brackets. This is called the roster method:

A={2,3,5,7}.A = \{2, 3, 5, 7\}.

A list may be too long to write out, or infinite, or of a length depending on some variable. In those cases we write out enough of it to make the pattern plain and leave the rest implicit, using an ellipsis:

N={1,2,3,…},{2,4,6,…,100}.\NN = \{1, 2, 3, \ldots\}, \qquad \{2, 4, 6, \ldots, 100\}.

Remark.

There is some genuine ambiguity here, since the reader is expected to work out what pattern the "…\ldots" is meant to suggest. Usually that is obvious. Sometimes it is not: {0,n,3,2,…}\{0, n, 3, 2, \ldots\} is meaningless.

In contrast to a list, the second notation describes the members implicitly by a rule. This is set-builder notation: we write

{ x∣x has some property },\{\, x \mid x \text{ has some property} \,\},

read “the set of all elements xx such that xx has that property”, the bar being read as “such that”. So the rationals, defined earlier in words, can be written

Q={ ab  |  a∈Z and b∈Z with b≠0 }.\QQ = \left\{\, \frac{a}{b} \;\middle|\; a \in \ZZ \text{ and } b \in \ZZ \text{ with } b \neq 0 \,\right\}.

Each notation has its own strength. A list says exactly what is in the set and nothing about why those things and not others; a rule says exactly why and leaves you to work out what actually satisfies it. The set {2,3,5,7}\{2, 3, 5, 7\} names its four members at a glance and hides that they are the primes below 1010.

Problem 1.5.

Write each of these sets both ways, as a list and by a rule: the even integers between −6-6 and 66; the integers whose square is 99; the real numbers whose square is −1-1.

Two Special Sets

The Empty Set

Calling a set a collection of objects is slightly misleading, because it suggests there has to be something there. There does not.

Every set carries with it a test for membership: given an object, either it belongs or it does not. Nothing in that test requires any object to pass it. Consider the set of all real numbers greater than 55 but less than 33. Any such number would belong to the set, and there is no such number, so the collection is empty, but it is still well defined, because we can still answer the membership question for every object we are handed.

Definition 1.11 (The Empty Set).

The set with no elements is called the empty set, written ∅\varnothing.

Remark.

A set is well defined when, given any object, there is an objective rule deciding whether the object belongs to it or not: one of the two answers must happen, and not both. The sets we study are always of this kind. “The set of all large numbers” is not a set in this sense, since nothing decides whether 10610^6 belongs to it.

Remark (The empty set is not zero).

The number 00 and the empty set are not the same thing. Take the set of all numbers that are neither positive nor negative. Exactly one number qualifies, so the set is {0}\{0\}, which has one element and so is not empty. A box with a zero in it is not an empty box.

The Universe of Discourse

At certain times we wish to limit membership in a set not only by the rule but also by what is eligible for membership in the first place: when solving an equation about lengths, nobody intends “the colour blue” to be a candidate.

Definition 1.12 (Universe of Discourse).

The universe of discourse, usually written UU, is a set fixed in advance that contains every eligible object under discussion, so that for every eligible object bb and every set AA under discussion, b∈U and A⊂U.b \in U \text{ and } A \subset U.

So, for instance, we may write { k∈{1,2,3,4}∣k⩾2 }={2,3,4},\{\, k \in \{1,2,3,4\} \mid k \geqslant 2 \,\} = \{2, 3, 4\}, where the eligible objects are the four listed and the rule then keeps three of them.

Once a universe is fixed, every set we can speak of is caught between two bounds: no set can have fewer elements than the empty set, and none can contain an object the universe does not, so

∅⊂A⊂U.\varnothing \subset A \subset U.
UAxy
Figure 1.9. A set AA drawn inside its universe UU, which is the rectangle surrounding it. Here x∈Ax \in A, while y∈Uy \in U but y∉Ay \notin A.

A picture of this kind, with sets drawn as regions and the universe as the rectangle around them, is called a Venn diagram. A Venn diagram helps us see an argument about sets, but it does not prove one.

The Arithmetic of Sets

Just as we combine two numbers to get a third, we can combine two sets to get a third. Doing so gives us an arithmetic of sets, and it has its own operations, its own rules, and its own closure.

Throughout this section every set is a subset of a fixed universe of discourse UU, as above.

Union

Definition 1.13 (Union).

The union of sets AA and BB, written A∪BA \cup B, is the set of all elements belonging to at least one of AA and BB:

A∪B={ x∣x∈A or x∈B }.A \cup B = \{\, x \mid x \in A \text{ or } x \in B \,\}.
UAB
Figure 1.10. The union A∪BA \cup B is the whole of the shaded region.

The word “or” here is the inclusive one: it means at least one, not exactly one. If Tom goes to the shop, or Jerry does, or both of them do, then in every one of those three cases somebody from the union went.

Intersection

Definition 1.14 (Intersection).

The intersection of sets AA and BB, written A∩BA \cap B, is the set of all elements belonging to both AA and BB at once:

A∩B={ x∣x∈A and x∈B }.A \cap B = \{\, x \mid x \in A \text{ and } x \in B \,\}.
UAB
Figure 1.11. The intersection A∩BA \cap B is the overlap.

The intersection of two sets can be empty even when neither set is. Take U=ZU = \ZZ, let AA be the even integers and BB the odd ones; no integer is both, so A∩B=∅A \cap B = \varnothing. This is one of the reasons the empty set has to exist at all: without it, the intersection of two perfectly good sets would sometimes fail to be a set.

Complement

Definition 1.15 (Complement).

The complement of a set AA, written A′A', is the set of all elements of the universe UU that are not in AA:

A′={ x∣x∈U but x∉A }.A' = \{\, x \mid x \in U \text{ but } x \notin A \,\}.
UAA′
Figure 1.12. The complement A′A' is everything inside the universe but outside AA.

The complement depends entirely on the universe. We never speak of all the non-AA’s in the world; only of all the elements of UU which are not in AA. Since everything under discussion lies in UU anyway, we could equally have written A′=U∩A′A' = U \cap A', and it is sometimes useful to remember that UU is there even when it is not written.

Relative Complement

Definition 1.16 (Relative Complement).

The relative complement of AA in BB, also called the difference and written B−AB - A, is the set of all elements of BB that are not in AA:

B−A={ x∣x∈B but x∉A }.B - A = \{\, x \mid x \in B \text{ but } x \notin A \,\}.
UAB
Figure 1.13. The difference B−AB - A: the part of BB that does not overlap AA.

The difference is not really a fourth operation, because it can be written with the three we already have:

B−A=B∩A′.B - A = B \cap A'.

An element of B−AB - A is one that is in BB and is not in AA, and “not in AA” is exactly “in A′A'”.

Combining the Operations

Expressions can be built up out of these operations as freely as arithmetic expressions are built out of ++ and ⋅\cdot, with brackets settling the order. So A∩(B∪C)A \cap (B \cup C) is read ”AA intersect the union of BB and CC”, and moving the bracket to give (A∩B)∪C(A \cap B) \cup C need not leave the set alone.

One combination splits a union into three pieces:

A∪B=(A∩B′)∪(A∩B)∪(B∩A′).A \cup B = (A \cap B') \cup (A \cap B) \cup (B \cap A').

The three pieces on the right are disjoint (no element lies in two of them), and they name the three regions of the diagram: the part of AA outside BB, the overlap, and the part of BB outside AA.

A ∩ BUA ∩ B′B ∩ A′
Figure 1.14. A union broken into three disjoint parts.

What is True of All Sets

Not everything that looks natural is true. It is tempting to expect complementation to pass through a union the way a minus sign passes through a bracket, giving (A∪B)′=A′∪B′(A \cup B)' = A' \cup B', and that is false in general. What is true is

(A∪B)′=A′∩B′,(A \cup B)' = A' \cap B',

which is one of De Morgan’s laws: to be outside both AA and BB is to be outside AA and also outside BB, and the union has turned into an intersection along the way.

What we want in this arithmetic are statements that hold for all sets, not ones that happen to come out right for a convenient choice of AA and BB.

Problem 1.6.

Draw the two-circle diagram and shade (A∪B)′(A \cup B)'. Then shade A′∪B′A' \cup B' on a fresh copy, and A′∩B′A' \cap B' on another. Which two agree?

Problem 1.7.

Take U={1,2,3,4,5,6}U = \{1, 2, 3, 4, 5, 6\}, A={1,2,3}A = \{1, 2, 3\} and B={3,4}B = \{3, 4\}. Write out A∪BA \cup B, A∩BA \cap B, A′A', B−AB - A and (A∪B)′(A \cup B)' as lists.

Problem 1.8.

Is A−BA - B the same set as B−AB - A? Compare with a−ba - b and b−ab - a for numbers.

Finally, both of the operations we started with are closed: given two sets, A∪BA \cup B is a set and A∩BA \cap B is a set, so the results can be combined again in their turn.

The Inclusion–Exclusion Principle

Write #(X)\#(X) for the number of elements of a finite set XX. The question of this section is what #\# does to a union, and the answer is not the one most people guess.

Counting a Union

A common mistake is to treat ∪\cup as though it were ++. A union does combine two sets, but it is not numerically additive unless the sets are completely separate.

Suppose 2020 students take Math 209, making up a set AA, and 3030 students take Bio 232, making up a set BB. Intuition says #(A∪B)=20+30=50\#(A \cup B) = 20 + 30 = 50, and that is right if and only if no student takes both courses. Suppose 77 students take both. Those 77 are among the 2020 and they are also among the 3030, so adding 20+3020 + 30 counts each of them twice, although each of them is one person. That leaves 1313 taking only Math 209 and 2323 taking only Bio 232.

UAB13237
Figure 1.15. The three regions hold 1313, 77 and 2323 students, so 4343 students in all.

Counting the diagram region by region gives 13+7+23=4313 + 7 + 23 = 43 distinct students, and the general rule follows.

For finite sets AA and BB,

#(A∪B)=#(A)+#(B)−#(A∩B),\#(A \cup B) = \#(A) + \#(B) - \#(A \cap B),

and this is called the inclusion–exclusion principle.

Here 20+30−7=4320 + 30 - 7 = 43, as counting the regions said it should be.

The intersection is what matters. Every element of A∩BA \cap B contributes 11 to #(A)\#(A) and another 11 to #(B)\#(B) although it is only one element, so the sum #(A)+#(B)\#(A) + \#(B) counts the intersection exactly twice, and subtracting it once puts things right. The wrong answer 5050 minus the correct answer 4343 is 77, which is the size of the intersection.

What the Two Sizes Alone Will Tell You

If all we are told is #(A)=20\#(A) = 20 and #(B)=30\#(B) = 30, we cannot recover #(A∪B)\#(A \cup B); we can only bound it, since

#(A∪B)=50−#(A∩B).\#(A \cup B) = 50 - \#(A \cap B).

At one extreme the sets are disjoint, #(A∩B)=0\#(A \cap B) = 0 and #(A∪B)=50\#(A \cup B) = 50. At the other, AA sits entirely inside BB, so #(A∩B)=20\#(A \cap B) = 20 and #(A∪B)=30\#(A \cup B) = 30. The union therefore has at least 3030 and at most 5050 elements, and by choosing the intersection we can arrange for any answer between the two.

Adding is Safe on Disjoint Regions

Behind all of this is one rule: numbers may be added when the regions they count are mutually exclusive. The three regions A∩B′A \cap B', A∩BA \cap B and A′∩BA' \cap B of Figure 1.14 do not overlap, so

#(A∪B)=#(A∩B′)+#(A∩B)+#(A′∩B),\#(A \cup B) = \#(A \cap B') + \#(A \cap B) + \#(A' \cap B),

and this addition needs no correction, because no element has been counted twice. The subtraction in the two-set formula is needed because #(A)\#(A) and #(B)\#(B) count overlapping regions.

Remark (Finite sets only).

From here on every set is finite. The formula becomes troublesome for infinite sets, because we cannot reliably subtract one infinity from another. Take A=NA = \NN. If BB is the set of even natural numbers then A−BA - B is the odd ones, which is infinite; but if BB is instead the set of natural numbers greater than 1010, then A−B={1,2,…,10}A - B = \{1, 2, \ldots, 10\}, which has exactly 1010 elements. In both cases AA and BB are infinite, so no arithmetic on their sizes alone can tell those two answers apart.

Example 1.17 (A Hand of Cards).

What is the chance of drawing a card that is either a heart or a face card from a standard deck of 5252?

There are 1313 hearts, making a set HH, and 1616 face cards (aces, kings, queens and jacks), making a set FF. The trap is to answer 13+16=2913 + 16 = 29. Four cards, the ace, king, queen and jack of hearts, lie in both sets, so #(H∩F)=4\#(H \cap F) = 4 and

#(H∪F)=13+16−4=25.\#(H \cup F) = 13 + 16 - 4 = 25.

So the chance is 25/5225/52, not the 29/5229/52 the intuitive count gives.

Three Sets

For finite sets AA, BB and CC the principle reads

#(A∪B∪C)=  #(A)+#(B)+#(C)−#(A∩B)−#(A∩C)−#(B∩C)+#(A∩B∩C).\begin{aligned} \#(A \cup B \cup C) = \;& \#(A) + \#(B) + \#(C) \\ & - \#(A \cap B) - \#(A \cap C) - \#(B \cap C) \\ & + \#(A \cap B \cap C). \end{aligned}

The name says what is happening: include, then exclude, then include again. Follow an element that lies in all three sets. The first line counts it three times, once in each of #(A)\#(A), #(B)\#(B) and #(C)\#(C). The second line subtracts it three times, once for each pair, which leaves it counted zero times. The third line adds it back once. It ends up counted once, which is correct. An element in exactly two of the sets is counted twice on the first line, subtracted once on the second and untouched by the third, so it too is counted once.

Example 1.18 (Language Classes).

Every student in a graduating class must take at least one of French, German or Spanish, giving sets FF, GG and SS with

#(F)=90,#(G)=80,#(S)=110,#(F∩G)=35,#(F∩S)=40,#(G∩S)=50,\begin{aligned} \#(F) &= 90, & \#(G) &= 80, & \#(S) &= 110, \\ \#(F \cap G) &= 35, & \#(F \cap S) &= 40, & \#(G \cap S) &= 50, \end{aligned}

and #(F∩G∩S)=20\#(F \cap G \cap S) = 20. How large is the class?

Straight from the formula,

#(F∪G∪S)=90+80+110−35−40−50+20=280−125+20=175.\#(F \cup G \cup S) = 90 + 80 + 110 - 35 - 40 - 50 + 20 = 280 - 125 + 20 = 175.

The diagram gets there another way, by filling the seven disjoint regions from the inside out. Put 2020 in the centre. Each pair intersection already contains that 2020, so the pair-only regions hold 35−20=1535 - 20 = 15, 40−20=2040 - 20 = 20 and 50−20=3050 - 20 = 30. Each single set contains the three regions just filled, so the parts taken by one language alone hold

90−(15+20+20)=35,80−(15+20+30)=15,110−(20+20+30)=40.90 - (15 + 20 + 20) = 35, \qquad 80 - (15 + 20 + 30) = 15, \qquad 110 - (20 + 20 + 30) = 40.
FGS35152015304020
Figure 1.16. The seven disjoint regions of F∪G∪SF \cup G \cup S, filled from the centre outwards.

Adding the seven regions, which is safe because they do not overlap,

35+15+40+15+20+30+20=175,35 + 15 + 40 + 15 + 20 + 30 + 20 = 175,

agreeing with the formula. The diagram carries more than the total: it also shows, for instance, that 1515 students took German and neither of the other two.

More Than Three

The pattern continues, with the signs alternating. Add the sets taken one at a time; subtract the intersections taken two at a time; add those taken three at a time; subtract those taken four at a time; and carry on alternating until the intersection of all of them has been used.

Problem 1.9.

In a group of 5050 people, 3030 read the news online and 2525 read it on paper. What is the largest the overlap could be, and what is the smallest, given that everybody reads it one way or the other?

Problem 1.10.

Write out the four-set formula in full. How many terms does it have, and how many of them carry a minus sign?

The Real Numbers

The first section of this chapter named the real numbers, as the points of the number line, and left it at that. The reason for leaving it there is that the real numbers are notoriously awkward to pin down, and doing it properly is the business of an analysis course rather than this one. For now the following will serve.

Definition 1.19 (Real Numbers, Informally).

A real number is a number that can be written as a decimal expansion, possibly one that never terminates and never repeats. These are exactly the numbers represented by the points of the number line: every point is a real number, and every real number is a point.

There also exist irrational numbers, meaning real numbers not contained in Q\QQ. One example is the ratio between the circumference of a circle and its diameter, which is known as π\pi; another is 2\sqrt{2}. The set of real numbers describes all physical quantities that can be represented on a line, and R\RR is the set of all rational and irrational numbers together.

The Rules of Arithmetic

The operations ++ and ⋅\cdot in R\RR — and so also in N\NN, Z\ZZ and Q\QQ — satisfy the following properties. Every manipulation in the rest of these notes is built out of them.

Let a,b,c∈Ra, b, c \in \RR.

  1. Commutativity. a+b=b+aa + b = b + a and a⋅b=b⋅aa \cdot b = b \cdot a.
  2. Associativity. (a+b)+c=a+(b+c)(a + b) + c = a + (b + c) and (a⋅b)⋅c=a⋅(b⋅c)(a \cdot b) \cdot c = a \cdot (b \cdot c).
  3. Distributivity. (a+b)⋅c=a⋅c+b⋅c(a + b) \cdot c = a \cdot c + b \cdot c and a⋅(b+c)=a⋅b+a⋅ca \cdot (b + c) = a \cdot b + a \cdot c.
  4. Neutral elements. a+0=0+a=aa + 0 = 0 + a = a and a⋅1=1⋅a=aa \cdot 1 = 1 \cdot a = a.
  5. Inverses. There is a number −a-a with a+(−a)=0a + (-a) = 0, and if a≠0a \neq 0 there is a number 1/a1/a with a⋅1/a=1a \cdot 1/a = 1.

We call 00 the neutral element with respect to ++ and 11 the neutral element with respect to ⋅\cdot, and we call −a-a the negative of aa and 1/a1/a the reciprocal of aa. None of these properties is claimed for subtraction or division, and in general several of them fail there; that is the point of Exercise 1.9.

What Follows from the Rules

Much of ordinary algebra follows from these five rules. We record the pieces we shall need.

If a+b=0a + b = 0 then b=−ab = -a and a=−ba = -b. Adding −a-a to both sides of a+b=0a + b = 0 gives −a+a+b=−a-a + a + b = -a, and the left-hand side is 0+b=b0 + b = b, so b=−ab = -a. Adding −b-b to both sides instead gives a=−ba = -b in the same way.

As a matter of convention, we shall write

a−binstead ofa+(−b).a - b \qquad\text{instead of}\qquad a + (-b).

With that convention, associativity and commutativity mean that a sum involving three terms may be written in many ways, all of them the same number:

(a−b)+c=(a+(−b))+c=a+(−b+c)=a+(c−b)=(a+c)−b,(a - b) + c = \bigl(a + (-b)\bigr) + c = a + (-b + c) = a + (c - b) = (a + c) - b,

so we may also write this sum as a−b+c=a+c−ba - b + c = a + c - b.

−(−a)=a-(-a) = a. By definition −a-a is the number that adds to aa to give 00; but (−a)+a=0(-a) + a = 0 says exactly that aa is the number which adds to −a-a to give 00, and that number is −(−a)-(-a).

−(a+b)=−a−b-(a + b) = -a - b. We show this by adding −a−b-a - b to a+ba + b: rearranging, (a+b)+(−a−b)=(a+(−a))+(b+(−b))=0+0=0(a + b) + (-a - b) = \bigl(a + (-a)\bigr) + \bigl(b + (-b)\bigr) = 0 + 0 = 0, so −a−b-a - b is the negative of a+ba + b.

0⋅a=00 \cdot a = 0. Here distributivity does the work:

0⋅a+a=0⋅a+1⋅a=(0+1)⋅a=1⋅a=a.0 \cdot a + a = 0 \cdot a + 1 \cdot a = (0 + 1) \cdot a = 1 \cdot a = a.

Adding −a-a to both sides gives 0⋅a+a−a=a−a=00 \cdot a + a - a = a - a = 0, and the left-hand side is simply 0⋅a+0=0⋅a0 \cdot a + 0 = 0 \cdot a. So 0⋅a=00 \cdot a = 0.

(−1)⋅a=−a(-1) \cdot a = -a. We show this the same way:

(−1)⋅a+a=(−1)⋅a+1⋅a=(−1+1)⋅a=0⋅a=0,(-1) \cdot a + a = (-1) \cdot a + 1 \cdot a = (-1 + 1) \cdot a = 0 \cdot a = 0,

and a number that adds to aa to give 00 is the negative of aa.

−(ab)=(−a)b-(ab) = (-a)b. What has to be shown is that (−a)b(-a)b is the negative of abab, which amounts to showing that the two add to zero; and by distributivity ab+(−a)b=(a+(−a))b=0⋅b=0ab + (-a)b = \bigl(a + (-a)\bigr)b = 0 \cdot b = 0.

−(ab)=a(−b)-(ab) = a(-b). We show this by the same computation with the factors the other way round: ab+a(−b)=a(b+(−b))=a⋅0=0ab + a(-b) = a\bigl(b + (-b)\bigr) = a \cdot 0 = 0.

(−a)(−b)=ab(-a)(-b) = ab. Applying the previous two facts in turn, (−a)(−b)=−(a(−b))=−(−(ab))=ab(-a)(-b) = -\bigl(a(-b)\bigr) = -\bigl(-(ab)\bigr) = ab.

Problem 1.11.

A number has only one negative. Show this: if a+b=0a + b = 0 and a+c=0a + c = 0, then b=cb = c.

Problem 1.12.

A nonzero number has only one reciprocal. Show this: if ab=1ab = 1 and ac=1ac = 1, then b=cb = c.

Problem 1.13.

Show that if ab=0ab = 0 then a=0a = 0 or b=0b = 0.

Fractions

With the rules of arithmetic written down we can say precisely how fractions behave, since every rule about them comes out of those five.

Throughout this section mm, nn, rr, ss are integers, and any letter standing in a denominator is nonzero.

The rule for cross-multiplying says that

mn=rsif and only ifms=rn.\frac{m}{n} = \frac{r}{s} \qquad\text{if and only if}\qquad ms = rn.

This is the rule that turns a question about two fractions into a question about two integers, and we use it whenever two fractions have to be compared.

The cancellation rule says that for a nonzero integer aa,

aman=mn.\frac{am}{an} = \frac{m}{n}.

To test the equality we apply the rule for cross-multiplying, so what has to hold is (am)n=m(an)(am)n = m(an), and that is true by associativity and commutativity. For instance,

45=(−2)(4)(−2)(5)=−8−10.\frac{4}{5} = \frac{(-2)(4)}{(-2)(5)} = \frac{-8}{-10}.

In dealing with quotients of integers which may be negative, it is useful to observe that

−mn=m−n,\frac{-m}{n} = \frac{m}{-n},

which cross-multiplying turns into (−m)(−n)=mn(-m)(-n) = mn, something we already know. Both are equal to −m/n-m/n, and for this reason we shall write

−mn=−mn=m−n-\frac{m}{n} = \frac{-m}{n} = \frac{m}{-n}

without worrying about which of the three is meant.

Definition 1.20 (Lowest Form).

A fraction r/sr/s of positive integers is in lowest form if rr and ss have no common divisor other than 11.

Any positive rational number has an expression as a fraction in lowest form. Start from any way of writing it as a quotient of positive integers m/nm/n. We know 11 is a common divisor of mm and nn, and any common divisor is at most equal to mm or nn, so among all common divisors there is a greatest one; call it dd and write m=drm = dr and n=dsn = ds with rr and ss positive integers. Cancelling dd,

mn=drds=rs,\frac{m}{n} = \frac{dr}{ds} = \frac{r}{s},

and rr and ss have no common divisor left, since one would have made dd larger.

Cancellation also explains addition. Given m/nm/n and r/sr/s we have seen that

mn=smsn,rs=nrns,\frac{m}{n} = \frac{sm}{sn}, \qquad \frac{r}{s} = \frac{nr}{ns},

so both now have the common denominator snsn, and adding is then adding numerators:

mn+rs=ms+nrns.\frac{m}{n} + \frac{r}{s} = \frac{ms + nr}{ns}.

Multiplication needs no such preparation, since

mn⋅rs=mrns.\frac{m}{n} \cdot \frac{r}{s} = \frac{mr}{ns}.

These are the formulas the closure argument earlier in the chapter used, and the two of them together are why Q\QQ is closed under all four operations while Z\ZZ is not.

Problem 1.14.

Put 84/12684/126 in lowest form, and check the answer by cross-multiplying against the original.

Rearranging a Relation

Suppose three numbers are related by

a+b=c.a + b = c.

Adding −b-b to both sides gives a=c−ba = c - b: the bb has crossed the equals sign and changed sign. That move, together with the corresponding one for multiplication, is most of elementary equation solving.

Example 1.21.

Solve 3x+4=193x + 4 = 19. Adding −4-4 to both sides gives 3x=153x = 15; multiplying both sides by 1/31/3 gives x=5x = 5.

Distributivity is what lets a bracket be expanded before the rearranging starts. Used twice,

(x+2)2=(x+2)(x+2)=x(x+2)+2(x+2)=x2+2x+2x+4=x2+4x+4.(x + 2)^2 = (x + 2)(x + 2) = x(x + 2) + 2(x + 2) = x^2 + 2x + 2x + 4 = x^2 + 4x + 4.

Equations, Identities and Inequalities

Consider the following four expressions.

(x+2)2=2x+7[1](x+2)2=x2+4x+4[2]x−2>1[3](x+2)2[4]\begin{aligned} (x + 2)^2 &= 2x + 7 & &[1] \\ (x + 2)^2 &= x^2 + 4x + 4 & &[2] \\ x - 2 &> 1 & &[3] \\ (x + 2)^2 & & &[4] \end{aligned}

By inspection we see that these four expressions are all different in nature, and by investigating each of them in turn we shall identify the differences.

Equations

Substituting 11 for xx in the left-hand side (LHS) and right-hand side (RHS) of [1][1] separately, we find

LHS=(1+2)2=32=9,RHS=2+7=9,\text{LHS} = (1 + 2)^2 = 3^2 = 9, \qquad \text{RHS} = 2 + 7 = 9,

so LHS == RHS when x=1x = 1. Substituting 22 for xx in the same way, however,

LHS=(2+2)2=16,RHS=4+7=11,\text{LHS} = (2 + 2)^2 = 16, \qquad \text{RHS} = 4 + 7 = 11,

so LHS ≠\neq RHS when x=2x = 2. Now rearranging the original expression gives

x2+4x+4=2x+7⟹x2+2x−3=0⟹(x+3)(x−1)=0,x^2 + 4x + 4 = 2x + 7 \quad\Longrightarrow\quad x^2 + 2x - 3 = 0 \quad\Longrightarrow\quad (x + 3)(x - 1) = 0,

from which LHS == RHS if and only if either x+3=0x + 3 = 0, that is x=−3x = -3, or x−1=0x - 1 = 0, that is x=1x = 1. The equality holds for no other value of xx.

Definition 1.22 (Equation).

An expression of this type, in which the two sides are equal only for a number of distinct values of the unknown quantity, is called an equation. The process of finding those values is called solving the equation.

Sets give us a tidy way of stating the answer. The problem “solve (x+2)2=2x+7(x + 2)^2 = 2x + 7” becomes “find the solution set of (x+2)2=2x+7(x + 2)^2 = 2x + 7”, and we write S={−3,1}.S = \{-3, 1\}.

Remark.

The same question can also be answered in set-builder notation, where S={ x∈R∣(x+2)2=2x+7 },S = \{\, x \in \RR \mid (x + 2)^2 = 2x + 7 \,\}, which is a statement of a different kind: the list tells us what the set contains but not the property its members share, while the rule tells us the property but not the members. Solving the equation is exactly the work of turning the second description into the first.

Recalling that two sets are equal when each is contained in the other, this tells us what it means for two equations to be the same problem: two equations are equivalent when they have the same solution set. Each rearrangement above replaced an equation by an equivalent one, and that is what entitled us to read the answer off the final line.

Problem 1.15.

Are x2=xx^2 = x and x=1x = 1 equivalent? Compare their solution sets.

Identities

Now take [2][2]. Substituting 11 for xx in both sides,

LHS=(1+2)2=9,RHS=12+4⋅1+4=9,\text{LHS} = (1 + 2)^2 = 9, \qquad \text{RHS} = 1^2 + 4 \cdot 1 + 4 = 9,

and substituting −1-1 for xx as before,

LHS=(−1+2)2=1,RHS=(−1)2+4(−1)+4=1.\text{LHS} = (-1 + 2)^2 = 1, \qquad \text{RHS} = (-1)^2 + 4(-1) + 4 = 1.

Whatever other numerical value we substitute for xx we find that LHS == RHS, so it appears that the two sides agree for all values of xx. Expanding the bracket, as we did above, says the same thing in algebra; the following geometric illustration says it a third way, without any algebra at all. Consider a square of side x+2x + 2 units.

(x + 2)2x2x2x2x2x22x2x4
Figure 1.17. One square, counted two ways.

Since these two squares are identical, their areas are identical. The left one has area (x+2)2(x + 2)^2, and the right one has been cut into four pieces of areas x2x^2, 2x2x, 2x2x and 44. Hence (x+2)2=x2+4x+4(x + 2)^2 = x^2 + 4x + 4 for all values of xx, and we say that (x+2)2(x+2)^2 is identical to x2+4x+4x^2 + 4x + 4.

Definition 1.23 (Identity).

A relationship whose two sides are equal for any value of the unknown quantity is called an identity. Using the symbol ≡\equiv for “is identical to”, the relationship above is written (x+2)2≡x2+4x+4.(x + 2)^2 \equiv x^2 + 4x + 4.

The two sides of an identity are two forms of the same expression. We shall use the identity symbol whenever we are dealing with an identity relationship, and we strongly recommend that the reader does the same.

Problem 1.16.

State which of the following are equations and which are identities.

  1. x2−9=(x−3)(x+3)x^2 - 9 = (x - 3)(x + 3)
  2. p2+2p−3=3−2p−p2p^2 + 2p - 3 = 3 - 2p - p^2
  3. y−1=1yy - 1 = \dfrac{1}{y}
  4. 2qq2−1=1q−1+1q+1\dfrac{2q}{q^2 - 1} = \dfrac{1}{q - 1} + \dfrac{1}{q + 1}

Inequalities

The third relationship, x−2>1x - 2 > 1, is obviously different from the first two. Reading from left to right, the symbol >> means “is greater than”, and << means “is less than”.

By inspection we see that for x−2x - 2 to have a value greater than one, xx must have a value greater than three:

if x−2>1,thenx>3.\text{if } x - 2 > 1, \quad\text{then}\quad x > 3.

Definition 1.24 (Inequality).

A relationship between two expressions using one of <<, >>, ⩽\leqslant, ⩾\geqslant is called an inequality.

Here the number line does the explaining. Consider a line as being made up of adjacent points; then all the real values the variable xx can take are represented by positions of points on that line, and the position of a point to the left of a second point corresponds to a value of xx less than the value at the second point:

−99<1,−10<−5,…-99 < 1, \qquad -10 < -5, \qquad \ldots

The values of xx given by the statement x>3x > 3 can then be represented by a section of this line.

12345x > 312345x ⩾ 3
Figure 1.18. A hollow circle leaves the endpoint out; a solid one takes it in.

From this we see that not all values of xx satisfy the inequality, but that there is an infinite set of values which do: the solution of an inequality is a range, or several ranges, of values of the variable involved. Note that x=3x = 3 is not included in the range, and this is what the hollow circle records. For x⩾3x \geqslant 3, which means xx is greater than or equal to 33, the value x=3x = 3 is included in the range, and the circle is filled in.

Problem 1.17.

Find the range of values of xx for which the following inequalities are true, and illustrate the range on a number line.

  1. x+1⩽−1x + 1 \leqslant -1
  2. 0⩾x−40 \geqslant x - 4
  3. 3<4−x3 < 4 - x

Expressions

The fourth of the four, (x+2)2(x + 2)^2, differs from the other three: it asserts nothing. It is a term, not a relationship, so there is nothing to solve and nothing to check. It has a value once xx is given, and that is all.

Powers

We have been writing x2x^2 and (x+2)2(x + 2)^2 without comment, on the strength of school algebra. We set the notation down properly, because we are about to extend it beyond whole-number exponents.

Definition 1.25 (Powers with Natural Exponents).

Let a∈Ra \in \RR and n∈Nn \in \NN. Then ”aa raised to the power of nn” is defined as

an=a⋅a⋯a⏟n occurrences of a,a^n = \underbrace{a \cdot a \cdots a}_{n \text{ occurrences of } a},

where aa is called the base and nn the exponent.

Power Rules

Let a,b∈Ra, b \in \RR and m,n∈Nm, n \in \NN. Then

an⋅am=an+m,an⋅bn=(a⋅b)n,(an)m=anm.a^n \cdot a^m = a^{n+m}, \qquad a^n \cdot b^n = (a \cdot b)^n, \qquad (a^n)^m = a^{nm}.

Each of the three is a matter of counting the factors. For the first, writing both powers out gives nn copies of aa followed by mm more, which is n+mn + m copies in all:

an⋅am=a⋯a⏟n copies⋅a⋯a⏟m copies=an+m.a^n \cdot a^m = \underbrace{a \cdots a}_{n \text{ copies}} \cdot \underbrace{a \cdots a}_{m \text{ copies}} = a^{n+m}.

For the second, nn copies of aa beside nn copies of bb can be paired off, one aa to one bb, by commutativity, giving nn copies of a⋅ba \cdot b:

an⋅bn=a⋯a⏟n copies⋅b⋯b⏟n copies=(ab)⋯(ab)⏟n copies=(ab)n.a^n \cdot b^n = \underbrace{a \cdots a}_{n \text{ copies}} \cdot \underbrace{b \cdots b}_{n \text{ copies}} = \underbrace{(ab) \cdots (ab)}_{n \text{ copies}} = (a b)^n.

For the third, (an)m(a^n)^m is mm copies of a block of nn copies of aa, which is nmnm copies of aa altogether:

(an)m=a⋯a⏟n⋯a⋯a⏟n⏟m blocks=anm.(a^n)^m = \underbrace{\underbrace{a \cdots a}_{n} \cdots \underbrace{a \cdots a}_{n}}_{m \text{ blocks}} = a^{nm}.

Remark.

We define a0=1a^0 = 1, and the reason is the first power rule: that is the only value making a0⋅an=a0+n=ana^0 \cdot a^n = a^{0+n} = a^n come out right for all n∈Nn \in \NN.

Definition 1.26 (Powers with Negative Integer Exponents).

Let a∈R∗a \in \RR^{*} and let n∈Nn \in \NN. Then ”aa raised to the power of −n-n” is defined as

a−n=1an.a^{-n} = \frac{1}{a^n}.

Remark.

Again the definition is forced rather than chosen. If the power rules are to hold for integer exponents then we need

a−n⋅an=a−n+n=a0=1,a^{-n} \cdot a^{n} = a^{-n+n} = a^0 = 1,

and this makes sense if and only if a−na^{-n} is the reciprocal of ana^n, which requires a≠0a \neq 0. With this definition the power rules above also hold for m,n∈Zm, n \in \ZZ.

Roots

How should the square root of a number aa be defined? Each of the obvious attempts has a problem.

  1. “The number which, raised to the power 22, gives aa.” But there may be two such numbers, since (−1)2=12=1(-1)^2 = 1^2 = 1.
  2. “A number which, raised to the power 22, gives aa.” Then −2-2 would be a square root of 44, and 4\sqrt{4} would not name one number.
  3. ”a1/2a^{1/2}.” But we have not yet said what raising to the power 1/21/2 means.

The way out is to demand a single number by insisting that it be the non-negative one.

Definition 1.27 (Square Root).

For a⩾0a \geqslant 0, the square root of aa, written a\sqrt{a} or a1/2a^{1/2}, is the non-negative real number such that (a)2=a\left(\sqrt{a}\right)^2 = a.

Remark.

Each part of that definition matters, in particular the condition a⩾0a \geqslant 0 and the word non-negative. Both 22 and −2-2 square to 44; only 22 is 4\sqrt{4}.

Definition 1.28 (nn-th Root).

Let n∈Nn \in \NN. For a⩾0a \geqslant 0, the nn-th root of aa, written an\sqrt[n]{a} or a1/na^{1/n}, is the non-negative real number such that (an)n=a\left(\sqrt[n]{a}\right)^n = a.

With roots in hand, an exponent may be any rational number at all.

Definition 1.29 (Powers with Rational Exponents).

For a>0a > 0 and a rational exponent r=p/qr = p/q with p∈Zp \in \ZZ and q∈Nq \in \NN, ”aa raised to the power of rr” is defined as

ar=(ap)1/q.a^{r} = \left(a^{p}\right)^{1/q}.

For a=0a=0, this definition is used only when p>0p>0; expressions with a zero base and a non-positive rational exponent are undefined.

The power rules above hold with mm and nn rational whenever all the expressions involved are defined; in particular, they may be used freely with positive bases.

Problem 1.18.

Simplify each of the following.

  1. y1/6⋅y−2/3y1/4\dfrac{y^{1/6} \cdot y^{-2/3}}{y^{1/4}}
  2. (t)3⋅t2t5\dfrac{\left(\sqrt{t}\right)^3 \cdot t^2}{\sqrt{t^5}}
  3. x2+x5/2x−1/2\dfrac{x^2 + x^{5/2}}{x^{-1/2}}
  4. m3/2−m−1/2m1/2+m−1/2\dfrac{m^{3/2} - m^{-1/2}}{m^{1/2} + m^{-1/2}}

Remarkable Identities

Three expansions come up so often that they are worth knowing by sight rather than working out each time.

Let a,b∈Ra, b \in \RR. Then

(a+b)2=a2+2ab+b2,(a−b)2=a2−2ab+b2,(a+b)(a−b)=a2−b2.(a + b)^2 = a^2 + 2ab + b^2, \qquad (a - b)^2 = a^2 - 2ab + b^2, \qquad (a + b)(a - b) = a^2 - b^2.

The first comes out of distributivity used twice, exactly as the expansion of (x+2)2(x + 2)^2 did:

(a+b)2=(a+b)(a+b)=a(a+b)+b(a+b)=a2+ab+ba+b2=a2+2ab+b2,(a + b)^2 = (a + b)(a + b) = a(a + b) + b(a + b) = a^2 + ab + ba + b^2 = a^2 + 2ab + b^2,

the last step by commutativity. The other two go the same way.

The first identity is also the picture we drew for (x+2)2(x + 2)^2, with the 22 replaced by bb: the big square has area (a+b)2(a+b)^2, and it decomposes into smaller rectangles of areas a2a^2, abab, abab and b2b^2.

ababa2ababb2
Figure 1.19. A visualisation of (a+b)2=a2+2ab+b2(a + b)^2 = a^2 + 2ab + b^2.

Problem 1.19.

Use the third identity to compute 101×99101 \times 99 in your head.

Surds

Expressions such as 4\sqrt{4} and 25\sqrt{25} have exact numerical values, namely 4=2\sqrt{4} = 2 and 25=5\sqrt{25} = 5. Expressions such as 2\sqrt{2}, 3\sqrt{3} and 5\sqrt{5} cannot be written as exact terminating or repeating decimals. The root form itself is exact. We might say that

2=1.4correct to 2 significant figures,\sqrt{2} = 1.4 \quad\text{correct to 2 significant figures,}

or that

2=1.4142136correct to 8 significant figures,\sqrt{2} = 1.4142136 \quad\text{correct to 8 significant figures,}

but we cannot write down an exact terminating or repeating decimal equal to 2\sqrt{2}. Such numbers are the irrational numbers met earlier, and it is often convenient to leave them in the exact forms 2\sqrt{2}, 3\sqrt{3}, and so on.

Definition 1.30 (Surd).

A root left in its root form, rather than replaced by a decimal approximation, is called a surd.

Remark.

Recall that 2\sqrt{2} means the positive square root of 22. So although x2=4x^2 = 4 is solved by both x=2x = 2 and x=−2x = -2, we still have 4=2\sqrt{4} = 2.

Surds occur frequently in solutions, so it is useful to be able to simplify them. The tool is the power rule ab=ab\sqrt{ab} = \sqrt{a}\sqrt{b}: look for a square factor and take it outside.

Example 1.31.

Express 48\sqrt{48} as the simplest possible surd.

48=16×3=16⋅3=43.\sqrt{48} = \sqrt{16 \times 3} = \sqrt{16} \cdot \sqrt{3} = 4\sqrt{3}.

Example 1.32.

Expand and simplify (2−33)(3+23)(2 - 3\sqrt{3})(3 + 2\sqrt{3}) and (5−27)(5+27)(5 - 2\sqrt{7})(5 + 2\sqrt{7}).

For the first,

(2−33)(3+23)=6−93+43−6(3)2=6−53−6×3=−12−53.\begin{aligned} (2 - 3\sqrt{3})(3 + 2\sqrt{3}) &= 6 - 9\sqrt{3} + 4\sqrt{3} - 6\left(\sqrt{3}\right)^2 \\ &= 6 - 5\sqrt{3} - 6 \times 3 \\ &= -12 - 5\sqrt{3}. \end{aligned}

For the second, the brackets are of the form (a+b)(a−b)(a + b)(a - b), so the middle terms cancel:

(5−27)(5+27)=25−107+107−4(7)2=25−28=−3.\begin{aligned} (5 - 2\sqrt{7})(5 + 2\sqrt{7}) &= 25 - 10\sqrt{7} + 10\sqrt{7} - 4\left(\sqrt{7}\right)^2 \\ &= 25 - 28 \\ &= -3. \end{aligned}

Problem 1.20.

Expand and simplify.

  1. (22+1)(2−2)(2\sqrt{2} + 1)(\sqrt{2} - 2)
  2. (33−2)(33+2)(3\sqrt{3} - 2)(3\sqrt{3} + 2)
  3. (26−3)2\left(2\sqrt{6} - 3\right)^2

Rationalising the Denominator

When the solution to a problem comes out containing surds, it is accepted practice to leave the answer in surd form, simplified as far as possible, unless an approximation has been asked for. Simplifying a fractional answer can often be managed by removing the surds from the denominator, and this process is called rationalising the denominator.

Example 1.33.

Rationalise the denominator of 32\dfrac{3}{\sqrt{2}}.

32=32⋅22=322.\frac{3}{\sqrt{2}} = \frac{3}{\sqrt{2}} \cdot \frac{\sqrt{2}}{\sqrt{2}} = \frac{3\sqrt{2}}{2}.

The expansion of (5−27)(5+27)(5 - 2\sqrt{7})(5 + 2\sqrt{7}) above shows how this works in general: a bracket of the form (a+b)(a−b)(a + b)(a - b) gives a2−b2a^2 - b^2, and squaring removes a square-root surd in aa or in bb. So a denominator a+ba + b is rationalised by multiplying above and below by a−ba - b.

Example 1.34.

Simplify 3−51+35\dfrac{3 - \sqrt{5}}{1 + 3\sqrt{5}}.

We multiply numerator and denominator by 1−351 - 3\sqrt{5}:

3−51+35=(3−5)(1−35)(1+35)(1−35)=48−10512−(35)2=48−105−44=24−55−22=55−2422.\begin{aligned} \frac{3 - \sqrt{5}}{1 + 3\sqrt{5}} &= \frac{(3 - \sqrt{5})(1 - 3\sqrt{5})}{(1 + 3\sqrt{5})(1 - 3\sqrt{5})} = \frac{48 - 10\sqrt{5}}{1^2 - \left(3\sqrt{5}\right)^2} \\ &= \frac{48 - 10\sqrt{5}}{-44} = \frac{24 - 5\sqrt{5}}{-22} = \frac{5\sqrt{5} - 24}{22}. \end{aligned}

Problem 1.21.

Rationalise the denominator.

  1. 13+25\dfrac{1}{3 + 2\sqrt{5}}
  2. 16−5\dfrac{1}{\sqrt{6} - \sqrt{5}}
  3. 12+1+12−1\dfrac{1}{\sqrt{2} + 1} + \dfrac{1}{\sqrt{2} - 1}

Even and Odd

The last idea of the chapter needs nothing new and will be used constantly: splitting the integers in two.

Definition 1.35 (Even and Odd).

An integer is even if it is twice an integer, and odd if it is one more than twice an integer. In set notation, the even integers and the odd integers are

E={ a∈Z∣a=2k for some k∈Z },O={ a∈Z∣a=2k+1 for some k∈Z }.E = \{\, a \in \ZZ \mid a = 2k \text{ for some } k \in \ZZ \,\}, \qquad O = \{\, a \in \ZZ \mid a = 2k + 1 \text{ for some } k \in \ZZ \,\}.

Every integer belongs to exactly one of EE and OO: dividing by 22 leaves a remainder of either 00 or 11, and no integer can be written in both forms.

The Parity of a Sum

Let aa and bb be positive integers.

  1. If aa is even and bb is even, then a+ba + b is even.
  2. If aa is even and bb is odd, then a+ba + b is odd.
  3. If aa is odd and bb is even, then a+ba + b is odd.
  4. If aa is odd and bb is odd, then a+ba + b is even.

We show the first. If aa and bb are even then a=2ka = 2k and b=2lb = 2l for some integers kk and ll, so by distributivity

a+b=2k+2l=2(k+l),a + b = 2k + 2l = 2(k + l),

and k+lk + l is an integer, so a+ba + b is twice an integer. The other three are left as problems below; each is the same computation with a +1+1 carried along.

The Parity of a Square

Let aa be a positive integer. If aa is even then a2a^2 is even, and if aa is odd then a2a^2 is odd.

If a=2ka = 2k then a2=(2k)2=4k2=2(2k2)a^2 = (2k)^2 = 4k^2 = 2(2k^2), which is even. If a=2k+1a = 2k + 1 then

a2=(2k+1)2=4k2+4k+1=2(2k2+2k)+1,a^2 = (2k + 1)^2 = 4k^2 + 4k + 1 = 2(2k^2 + 2k) + 1,

which is one more than twice an integer, and so odd.

This can also be read backwards: if a2a^2 is even then aa is even. Every integer is either even or odd, and not both. If aa were odd, what we have just shown would make a2a^2 odd; but a2a^2 is even, so the odd case is ruled out, and aa must be even. Exercise 1.6 uses this step.

Problem 1.22.

Show the three remaining parts of the parity-of-a-sum result.

Problem 1.23.

Show that if nn is even then (−1)n=1(-1)^n = 1, and that if nn is odd then (−1)n=−1(-1)^n = -1.

Problem 1.24.

Show that if mm and nn are odd then the product mnmn is odd.

Exercises

Exercise 1.1.

Compute

(43+52)⋅65−25⋅52.\left(\frac{4}{3} + \frac{5}{2}\right) \cdot \frac{6}{5} - \frac{2}{5} \cdot \frac{5}{2}.

Exercise 1.2.

Simplify:

A=2+89+519,B=1+18+34−11256+13+2,C=51273⋅815−56.A = 2 + \cfrac{8}{9 + \cfrac{5}{19}}, \qquad B = \frac{1 + \frac{1}{8} + \frac{3}{4} - \frac{1}{12}}{\frac{5}{6} + \frac{1}{3} + 2}, \qquad C = \frac{\frac{5}{12}}{\frac{7}{3} \cdot \frac{8}{15} - \frac{5}{6}}.

Exercise 1.3.

Bigger, smaller, or equal?

(i) 514 and 621,(ii) 275 and 163,(iii) 313 and 2191.\text{(i) } \frac{5}{14} \text{ and } \frac{6}{21}, \qquad \text{(ii) } \frac{27}{5} \text{ and } \frac{16}{3}, \qquad \text{(iii) } \frac{3}{13} \text{ and } \frac{21}{91}.

Exercise 1.4.

Compute

(i) 1(234),(ii) 1(234),(iii) 1234.\text{(i) } \frac{1}{\left(\dfrac{2}{\frac{3}{4}}\right)}, \qquad \text{(ii) } \frac{1}{\left(\dfrac{\frac{2}{3}}{4}\right)}, \qquad \text{(iii) } \frac{\frac{1}{2}}{\frac{3}{4}}.

Exercise 1.5.

Let x,y,z∈Rx, y, z \in \RR be such that the following formulas are defined. Simplify:

(i) (4xy+3yz)(4zxy−2y),(ii) xy−zx2x+zy,(iii) x1−11−x.\text{(i) } \left(\frac{4}{xy} + \frac{3}{yz}\right)\left(\frac{4z}{xy} - \frac{2}{y}\right), \qquad \text{(ii) } \frac{\frac{x}{y} - z}{\frac{x^2}{x} + \frac{z}{y}}, \qquad \text{(iii) } \frac{x}{1 - \frac{1}{1 - x}}.

Exercise 1.6.

Let us assume 2\sqrt{2} is a rational number, that is, that there exist a,b∈Za, b \in \ZZ with a/b=2a/b = \sqrt{2}. Suppose this fraction is reduced as much as possible.

  1. By squaring the equation, show that aa is even.
  2. Is bb even or odd?
  3. Show that the two answers cannot both hold.
  4. What do you conclude about 2\sqrt{2}?

Exercise 1.7.

Taking inspiration from the previous exercise, show that p\sqrt{p} is not a rational number when pp is prime. (Reminder: an integer is called prime when it has exactly two divisors.)

Exercise 1.8.

Let x∈Qx \in \QQ and y∉Qy \notin \QQ. Show that x+y∉Qx + y \notin \QQ.

Exercise 1.9.

Show that division in R\RR is neither commutative, nor associative, nor distributive over ++.

Exercise 1.10.

Let a,b∈R+∗a, b \in \RR_{+}^{*}.

  1. If a=ba = b, can we say a=b\sqrt{a} = \sqrt{b}?
  2. If a=b\sqrt{a} = \sqrt{b}, can we say a=ba = b?
  3. Can we say a+b=a+b\sqrt{a + b} = \sqrt{a} + \sqrt{b}?

Exercise 1.11.

Show that 1+23=13+431 + 2\sqrt{3} = \sqrt{13 + 4\sqrt{3}}.

Exercise 1.12.

Let a,b∈R∗a, b \in \RR^{*}. Simplify the following expressions, where that is possible:

(i) (a−2b3)2,(ii) a2+aba3,(iii) a2+b2ab.\text{(i) } \left(a^{-2} b^3\right)^2, \qquad \text{(ii) } \frac{a^2 + ab}{a^3}, \qquad \text{(iii) } \frac{a^2 + b^2}{ab}.

Exercise 1.13.

Let a,b,c∈Ra, b, c \in \RR. What is (a+b+c)2(a + b + c)^2?

Exercise 1.14.

Let a>0a > 0 and b>0b > 0 with a≠ba \neq b. Show that

1a−b=a+ba−b.\frac{1}{\sqrt{a} - \sqrt{b}} = \frac{\sqrt{a} + \sqrt{b}}{a - b}.

Exercise 1.15.

Compute

6−2⋅6+2.\sqrt{\sqrt{6} - \sqrt{2}} \cdot \sqrt{\sqrt{6} + \sqrt{2}}.

Exercise 1.16.

Simplify:

(i) (−1+a)2−(1−a)2,(ii) (a2+b2)2−(a2−b2)2.\text{(i) } (-1 + a)^2 - (1 - a)^2, \qquad \text{(ii) } \left(a^2 + b^2\right)^2 - \left(a^2 - b^2\right)^2.

Exercise 1.17.

Let a,b,x>0a,b,x>0 and n∈Zn\in\ZZ. Simplify:

(i) 3an+1⋅6xn+7⋅9bn+13xn⋅2bn+1⋅3a,(ii) a33⋅a8.\text{(i) } \frac{3a^{n+1} \cdot 6x^{n+7} \cdot 9b^{n+1}}{3x^{n} \cdot 2b^{n+1} \cdot 3a}, \qquad \text{(ii) } \sqrt[3]{a^3} \cdot \sqrt{a^8}.

Exercise 1.18.

Let AA be the set whose elements are aa, bb, cc and dd.

  1. Write AA in roster form.
  2. How many subsets does AA have? How many of them have exactly two elements? List those by the roster method.
  3. If a fair coin is tossed four times, is there a fifty-fifty chance of getting two heads and two tails? How is that question related to the previous part?
  4. Now let AA have six members instead. How many subsets does it have, and how many of them have exactly two elements, exactly three, exactly four? Two of those three counts agree; say why.

Exercise 1.19.

The universe of discourse decides what an answer looks like.

  1. Write the solution set of x4=1x^4 = 1 in set-builder notation, and then in roster form when the universe is (i) the complex numbers, (ii) the real numbers, (iii) the positive real numbers, (iv) the even integers. (Reminder: the complex numbers are the numbers a+b ia + b\,i with aa and bb real and i2=−1i^2 = -1.)
  2. Write the solution set of (2x−3)(x+1)(x+7)=0(2x - 3)(x + 1)(x + 7) = 0 in set-builder notation, and then in roster form when the universe is (i) the rational numbers, (ii) the real numbers, (iii) the integers, (iv) the positive integers.

Exercise 1.20.

Brackets matter in set expressions just as they do in arithmetic.

  1. Show that (A∩B)∪C(A \cap B) \cup C and A∩(B∪C)A \cap (B \cup C) need not be equal. What is the most general condition on AA, BB and CC under which they are equal?
  2. Why is it ambiguous to write A∩B∪CA \cap B \cup C? Is it ambiguous to write A∩B∩CA \cap B \cap C?
  3. Show that (A∪B)∩C(A \cup B) \cap C and A∪(B∩C)A \cup (B \cap C) need not be equal, and that (A∪B)∩C(A \cup B) \cap C and (A∩C)∪(B∩C)(A \cap C) \cup (B \cap C) always are.

Exercise 1.21.

Cancelling is not as free for sets as it is for numbers.

  1. Show that it is possible to have X∪B=X∪CX \cup B = X \cup C and yet B≠CB \neq C.
  2. Show that if X∪B=X∪CX \cup B = X \cup C and also X∩B=X∩CX \cap B = X \cap C, then B=CB = C.

Exercise 1.22.

For real numbers aa and bb we say aa is less than bb, and write a<ba < b, if and only if b−ab - a is positive, that is, if b=a+hb = a + h for some positive hh. If a<ba < b we also say bb is greater than aa and write b>ab > a. Using only this definition and the rules of arithmetic, show the following.

  1. If a<ba < b then a+c<b+ca + c < b + c, and also a−c<b−ca - c < b - c.
  2. If a<ba < b and c<dc < d then a+c<b+da + c < b + d. Is a−c<b−da - c < b - d also true in this case? Explain.
  3. If a<ba < b and cc is positive then ac<bcac < bc. What happens if the restriction that cc be positive is dropped?
  4. If a<ba < b and aa and bb are either both positive or both negative, which is to say ab>0ab > 0, then 1/b<1/a1/b < 1/a.

Check Yourself

 

Fresh questions on the whole chapter — none of them is worked out above. Do each on paper first; the box only tells you whether you got there.

Answers are checked in your browser, as often as you like. Nothing is sent anywhere and nothing is kept but your own work. A formula may be written with the symbols themselves or with ~ & | -> <-> ^, and \and, \or, \to expand as you type.

Exercise 1.23.

Which one operation is the set {−1,0,1}\{-1, 0, 1\} closed under?

answer one of these

Exercise 1.24.

Exactly one of the following is true. Which?

answer one of these

Exercise 1.25.

For which value of aa does a0\dfrac{a}{0} fail because too many numbers qualify, rather than because none does?

answer one of these

Exercise 1.26.

Which of these sets is empty?

answer one of these

Exercise 1.27.

The chapter gives one of De Morgan’s two laws. What is (A∩B)′(A \cap B)' equal to, for all sets AA and BB?

answer one of these

Exercise 1.28.

For finite sets, #(A∩B′)\#(A \cap B') is equal to which of these?

answer one of these

Exercise 1.29.

Which of these is an identity rather than an equation?

answer one of these

Exercise 1.30.

What is the solution set of x2−x−6=0x^2 - x - 6 = 0?

answer one of these

Exercise 1.31.

For p>0p>0, simplify p1/2⋅p−3/4p−1/4\dfrac{p^{1/2} \cdot p^{-3/4}}{p^{-1/4}}.

answer one of these

Exercise 1.32.

Express 72\sqrt{72} as the simplest possible surd.

answer one of these

Exercise 1.33.

Expand and simplify (4−32)(4+32)(4 - 3\sqrt{2})(4 + 3\sqrt{2}).

answer one of these

Exercise 1.34.

Rationalise the denominator of 53\dfrac{5}{\sqrt{3}}.

answer one of these

Exercise 1.35.

Work out the value of

5+45−4.\frac{\sqrt{5} + \sqrt{4}}{\sqrt{5} - \sqrt{4}}.
answer one of these

Exercise 1.36.

A set AA has 1212 elements and a set BB has 1818. Which values can #(A∪B)\#(A \cup B) take?

answer one of these

Exercise 1.37.

Of 4040 students, 2222 play football and 1818 play chess, and 77 do both. How many play at least one of the two?

answer one of these

Exercise 1.38.

Let aa be a positive integer whose square a2a^2 is odd. What follows about aa?

answer one of these

Lesson 2

Functions, and Equations

Taught

Everything here rests on the arithmetic and algebra of the first lesson: the rules of arithmetic, powers and roots, surds, and the difference between an equation and an identity. What is new is the idea of a function, and the algebra that idea makes possible.

Functions

Definition 2.1 (Function).

A function, defined for all numbers, is an association which to each number associates another number. If we denote the function by ff, the association is written

f:x↦f(x),f : x \mapsto f(x),

and f(x)f(x) is called the value of the function at xx, or the image of xx under ff.

Example 2.2.

The association f:x↦x2f : x \mapsto x^2 is a function, called the square. The association g:x↦x+1g : x \mapsto x + 1 is another. Their values are read off by substituting:

f(2)=4,f(3)=9,f(−1)=1,f(−2)=4,g(1)=2,g(2)=3,g(50)=51.f(2) = 4, \quad f(3) = 9, \quad f(-1) = 1, \quad f(-2) = 4, \qquad g(1) = 2, \quad g(2) = 3, \quad g(50) = 51 .

Thinking of the value given to xx as an input and the corresponding value of the function as an output, we may read f:r↦2r+1f : r \mapsto 2r + 1 as “the function which, when we input a value for rr, gives as output the value 2r+12r + 1“.

fxf(x)inputoutputone output for each input
Figure 2.1. A function as a rule turning an input into an output. What makes it a function is that each input has exactly one output.

Using ff as the symbol for a function, we write f(x)f(x) for a function of xx and f(r)f(r) for a function of rr, so the two associations above may equally be written f(x)=(x+2)2f(x) = (x+2)^2 and f(r)=2r+1f(r) = 2r + 1. To represent the value when x=1x = 1 we write f(1)f(1). So for f:x↦(x+2)2f : x \mapsto (x+2)^2,

f(1)=(1+2)2=9,f(3)=(3+2)2=25,f(−2)=(−2+2)2=0.f(1) = (1+2)^2 = 9, \qquad f(3) = (3+2)^2 = 25, \qquad f(-2) = (-2+2)^2 = 0 .

Both xx and f(x)f(x) vary, but xx may be given any value while the value of f(x)f(x) depends on it. So xx is called the independent variable and f(x)f(x) the dependent variable.

Example 2.3.

The association which to each number xx associates the number 44 is the constant function with value 44. More generally, for a given number cc, the association

x↦cfor all numbers xx \mapsto c \quad \text{for all numbers } x

is the constant function with value cc.

Remark (A warning about language).

It is convenient, and slightly incorrect, to write a sentence like “let f(x)f(x) be such and such a function”. The trouble is that xx is not quantified: what is meant is the function whose value at a number xx is such and such. We shall try to avoid the loose phrasing at least while the idea is new.

The examples so far have been given by formulas. Functions may be defined quite arbitrarily, and no formula is required: to describe a function amounts to giving its values at all numbers for which it is defined.

Remark (What the values are).

We have adopted the convention that the values of a function are numbers. It is not a universal convention, and it is a useful one here.

The Algebra of Functions

Functions defined on the same set can be added and multiplied, and when they are they obey the rules of arithmetic from the first lesson.

Let ff and gg be functions defined on the same set SS. Their sum f+gf + g is the function whose value at an element xx of SS is

(f+g)(x)=f(x)+g(x).(f+g)(x) = f(x) + g(x).

Associativity of addition for numbers gives associativity for functions at once: for any ff, gg, hh defined on SS,

(f+g)+h=f+(g+h),(f + g) + h = f + (g + h),

and commutativity for numbers gives f+g=g+ff + g = g + f in the same way.

There is a zero function, whose value at every xx in SS is 00, and which we also denote by 00. For every ff defined on SS,

f+0=0+f=f.f + 0 = 0 + f = f .

If ff is defined on SS then minus ff, written −f-f, is the function whose value at xx is −f(x)-f(x).

Example 2.4.

If f(x)=x2f(x) = x^2 then (−f)(x)=−x2(-f)(x) = -x^2, so (−f)(5)=−25(-f)(5) = -25, and

f+(−f)=0,f + (-f) = 0,

the zero function.

So functions satisfy the same basic rules for addition as numbers do, and the same holds for multiplication. The product fgfg of two functions defined on SS is the function whose value at xx is

(fg)(x)=f(x) g(x),(fg)(x) = f(x)\,g(x),

and this product is commutative and associative for the same reason as before. If 11 denotes the constant function with value 11 for all xx in SS, then

1f=fand0f=0.1f = f \qquad\text{and}\qquad 0f = 0 .

Multiplication distributes over addition. Reading the definitions off one after another,

((f+g)h)(x)=(f+g)(x)⋅h(x)=(f(x)+g(x))h(x)=f(x)h(x)+g(x)h(x)=(fh)(x)+(gh)(x)=(fh+gh)(x),\begin{aligned} \bigl((f+g)h\bigr)(x) &= (f+g)(x) \cdot h(x) \\ &= \bigl(f(x) + g(x)\bigr)h(x) \\ &= f(x)h(x) + g(x)h(x) \\ &= (fh)(x) + (gh)(x) \\ &= (fh + gh)(x), \end{aligned}

and since the two sides agree at every xx in SS,

(f+g)h=fh+gh.(f+g)h = fh + gh .

Functions of numbers occur in physical life whenever one quantity is described in terms of another.

Problem 2.1.

Let f(x)=x2f(x) = x^2 and g(x)=x+1g(x) = x + 1, both defined for all numbers.

  1. Give the values of (f+g)(3)(f+g)(3), (fg)(3)(fg)(3) and (−f)(3)(-f)(3).
  2. Write down formulas for (f+g)(x)(f+g)(x) and (fg)(x)(fg)(x).
  3. Show that fgfg and gfgf have the same value at every xx.

Problem 2.2.

Find f(0)f(0), f(1)f(1), f(−2)f(-2) and f(5)f(5) for each of the following.

  1. f:x↦x2−3xf : x \mapsto x^2 - 3x;
  2. f:x↦(x−7)(x+2)f : x \mapsto (x-7)(x+2);
  3. f:x↦1x2+1f : x \mapsto \dfrac{1}{x^2 + 1}.

Polynomials

Definition 2.5 (Polynomial).

If every term of a function has the form axnax^n, where aa is a constant and nn is a non-negative integer, the function is a polynomial. The highest power of xx occurring in it is its degree.

Example 2.6.

The functions x2x^2, 2x2x, 3x5−7x+63x^5 - 7x + 6 and (x−4)2(x-4)^2 are polynomials; x\sqrt{x}, x−2\sqrt{x-2} and 1/x1/x are not, since their exponents are not non-negative integers (Definition 1.29 gives the meaning of x1/2x^{1/2}, and it is not a whole power).

The polynomial 5x6−7x3+6x5x^6 - 7x^3 + 6x has degree 66.

Fractional Functions

A fractional function is one of the form

p(x)q(x)\frac{p(x)}{q(x)}

with pp and qq polynomials, where qq is not the zero polynomial. Its domain consists of the real numbers for which q(x)≠0q(x)\neq 0. It is called proper if the degree of the numerator is less than the degree of the denominator, and improper if the degree of the numerator is greater than or equal to that of the denominator.

An improper numerical fraction such as 97\tfrac{9}{7} may be written as a whole number plus a proper fraction,

97=7+27=1+27,\frac{9}{7} = \frac{7+2}{7} = 1 + \frac{2}{7},

and an improper algebraic fraction splits the same way. The fraction x2−1x2+1\dfrac{x^2-1}{x^2+1} is improper, its numerator and denominator having equal degree, and

x2−1x2+1=(x2+1)−2x2+1=1−2x2+1.\frac{x^2 - 1}{x^2 + 1} = \frac{(x^2 + 1) - 2}{x^2 + 1} = 1 - \frac{2}{x^2+1}.

Problem 2.3.

State the degree of each polynomial, and say which of the following fractional functions are proper and which improper.

  1. x+3(x−2)(x+4)\dfrac{x+3}{(x-2)(x+4)};
  2. x3x2−1\dfrac{x^3}{x^2 - 1};
  3. 2x2−1x2+x+1\dfrac{2x^2 - 1}{x^2 + x + 1}.

Partial Fractions

Two fractions with different denominators may be combined into one. Starting from

f(x)=2x+1+xx2+1,f(x) = \frac{2}{x+1} + \frac{x}{x^2+1},

putting both over a common denominator gives

f(x)=2(x2+1)+x(x+1)(x+1)(x2+1)=3x2+x+2(x+1)(x2+1).f(x) = \frac{2(x^2+1) + x(x+1)}{(x+1)(x^2+1)} = \frac{3x^2 + x + 2}{(x+1)(x^2+1)} .

It is often useful to run this backwards: to take a fractional function and write it as a sum of separate fractions with simpler denominators. This is called decomposing the function into partial fractions. If the original fraction is proper then the partial fractions are proper too.

The shape of the decomposition is fixed by the factors of the denominator, and the constants in it are found by substitution or by comparing coefficients.

Linear Factors

A linear denominator carries a single constant on top.

Example 2.7 (A proper fraction with linear factors).

Decompose x+3(x−2)(x+4)\dfrac{x+3}{(x-2)(x+4)} into partial fractions.

The factors are linear, so we look for constants AA and BB with

x+3(x−2)(x+4)=Ax−2+Bx+4=A(x+4)+B(x−2)(x−2)(x+4).\frac{x+3}{(x-2)(x+4)} = \frac{A}{x-2} + \frac{B}{x+4} = \frac{A(x+4) + B(x-2)}{(x-2)(x+4)} .

The denominators are identical, so the numerators must be as well:

x+3=A(x+4)+B(x−2),x + 3 = A(x+4) + B(x-2),

and this holds for every value of xx. Substituting x=2x = 2 kills the BB term and gives 5=6A5 = 6A, so A=56A = \tfrac{5}{6}. Substituting x=−4x = -4 kills the AA term and gives −1=−6B-1 = -6B, so B=16B = \tfrac{1}{6}. Therefore

x+3(x−2)(x+4)=56(x−2)+16(x+4).\frac{x+3}{(x-2)(x+4)} = \frac{5}{6(x-2)} + \frac{1}{6(x+4)} .

A Quadratic Factor

A quadratic denominator that does not factorise carries a linear numerator, so that the partial fraction is still proper.

Example 2.8 (A quadratic factor in the denominator).

Express x−3(x−1)(x2+1)\dfrac{x-3}{(x-1)(x^2+1)} in partial fractions.

The second factor is quadratic, so we look for

x−3(x−1)(x2+1)=Ax−1+Bx+Cx2+1,\frac{x-3}{(x-1)(x^2+1)} = \frac{A}{x-1} + \frac{Bx+C}{x^2+1},

each numerator chosen so that its fraction is proper. Clearing denominators,

x−3=A(x2+1)+(Bx+C)(x−1).(∗)x - 3 = A(x^2+1) + (Bx+C)(x-1). \tag{$*$}

Substituting x=1x = 1 gives −2=2A-2 = 2A, so A=−1A = -1. There is no value of xx making x2+1x^2 + 1 zero, so AA cannot be eliminated the same way, and the remaining constants come from comparing coefficients. On the left the coefficient of x2x^2 is 00; on the right it is A+BA + B. Hence B=−A=1B = -A = 1. Substituting x=0x = 0 in (∗)(*) gives −3=A−C=−1−C-3 = A - C = -1 - C, so C=2C = 2. Therefore

x−3(x−1)(x2+1)=−1x−1+x+2x2+1.\frac{x-3}{(x-1)(x^2+1)} = \frac{-1}{x-1} + \frac{x+2}{x^2+1} .

Either route reaches the constants: substituting values chosen to eliminate terms, or comparing the coefficients of each power. In practice a mixture of the two is quickest.

Repeated Factors

Example 2.9 (A repeated factor).

Express x−1(x+1)(x−2)2\dfrac{x-1}{(x+1)(x-2)^2} in partial fractions.

Since (x−2)2(x-2)^2 is a repeated factor we might think of writing Bx+C(x−2)2\dfrac{Bx+C}{(x-2)^2} over it. That is not the simplest form. Putting C=−2B+DC = -2B + D,

Bx+C(x−2)2=Bx−2B+D(x−2)2=B(x−2)+D(x−2)2=Bx−2+D(x−2)2,\frac{Bx + C}{(x-2)^2} = \frac{Bx - 2B + D}{(x-2)^2} = \frac{B(x-2) + D}{(x-2)^2} = \frac{B}{x-2} + \frac{D}{(x-2)^2},

so a repeated linear factor is better written as two fractions. We therefore look for

x−1(x+1)(x−2)2=Ax+1+Bx−2+D(x−2)2,\frac{x-1}{(x+1)(x-2)^2} = \frac{A}{x+1} + \frac{B}{x-2} + \frac{D}{(x-2)^2},

that is, x−1=A(x−2)2+B(x+1)(x−2)+D(x+1)x - 1 = A(x-2)^2 + B(x+1)(x-2) + D(x+1). Substituting x=2x = 2 gives 1=3D1 = 3D, so D=13D = \tfrac13. Substituting x=−1x = -1 gives −2=9A-2 = 9A, so A=−29A = -\tfrac29. Comparing the coefficients of x2x^2 gives 0=A+B0 = A + B, so B=29B = \tfrac29. Therefore

x−1(x+1)(x−2)2=−29(x+1)+29(x−2)+13(x−2)2.\frac{x-1}{(x+1)(x-2)^2} = -\frac{2}{9(x+1)} + \frac{2}{9(x-2)} + \frac{1}{3(x-2)^2} .

In general a repeated factor (ax+b)2(ax+b)^2 in a denominator gives rise to two partial fractions,

Aax+bandB(ax+b)2,\frac{A}{ax+b} \qquad\text{and}\qquad \frac{B}{(ax+b)^2},

and a factor (ax+b)3(ax+b)^3 gives rise to three, with denominators ax+bax+b, (ax+b)2(ax+b)^2 and (ax+b)3(ax+b)^3.

Improper Fractions

An improper fraction has to be divided out first, leaving a whole part and a proper remainder.

Example 2.10 (An improper fraction).

Express x2+1(x+1)(x−3)\dfrac{x^2+1}{(x+1)(x-3)} in partial fractions.

The numerator and the denominator both have degree 22, so the fraction is improper. Since (x+1)(x−3)=x2−2x−3(x+1)(x-3) = x^2 - 2x - 3,

x2+1x2−2x−3=(x2−2x−3)+(2x+4)x2−2x−3=1+2x+4(x+1)(x−3),\frac{x^2+1}{x^2-2x-3} = \frac{(x^2 - 2x - 3) + (2x + 4)}{x^2 - 2x - 3} = 1 + \frac{2x+4}{(x+1)(x-3)},

and the remaining fraction is proper. Writing

2x+4(x+1)(x−3)=Ax+1+Bx−3,that is2x+4=A(x−3)+B(x+1),\frac{2x+4}{(x+1)(x-3)} = \frac{A}{x+1} + \frac{B}{x-3}, \qquad\text{that is}\qquad 2x + 4 = A(x-3) + B(x+1),

and substituting x=−1x = -1 gives 2=−4A2 = -4A, so A=−12A = -\tfrac12; substituting x=3x = 3 gives 10=4B10 = 4B, so B=52B = \tfrac52. Therefore

x2+1(x+1)(x−3)=1−12(x+1)+52(x−3).\frac{x^2+1}{(x+1)(x-3)} = 1 - \frac{1}{2(x+1)} + \frac{5}{2(x-3)} .

Problem 2.4.

Express each of the following in partial fractions.

  1. 5(x−1)(x+4)\dfrac{5}{(x-1)(x+4)};
  2. x(x−1)(x+1)2\dfrac{x}{(x-1)(x+1)^2};
  3. x2(x−1)(x−2)\dfrac{x^2}{(x-1)(x-2)}.

Quadratic Equations

Definition 2.11 (Quadratic Equation).

An equation of the form

ax2+bx+c=0,a≠0,ax^2 + bx + c = 0, \qquad a \neq 0,

with a,b,ca, b, c real numbers, is a quadratic equation. Its solutions are called its roots.

Solving by Factorising

If the left-hand side factorises, the equation is solved by setting each factor to zero, since a product of two numbers is zero only when one of them is.

Example 2.12.

Solve 2x2−7x+3=02x^2 - 7x + 3 = 0.

The left-hand side factorises, so the equation becomes

(2x−1)(x−3)=0,(2x-1)(x-3) = 0,

from which either 2x−1=02x - 1 = 0, giving x=12x = \tfrac12, or x−3=0x - 3 = 0, giving x=3x = 3.

Losing a Solution

Solutions can be lost if a step is taken carelessly. Consider 2x2−14x=02x^2 - 14x = 0 and two ways of handling it.

Dividing through by 2x2x gives x−7=0x - 7 = 0, so x=7x = 7.

Factorising instead gives 2x(x−7)=02x(x-7) = 0, so either 2x=02x = 0 or x−7=0x - 7 = 0, giving x=0x = 0 or x=7x = 7.

The first route lost the solution x=0x = 0, and it lost it because the equation was divided by the common factor xx. Dividing by a constant factor is correct and desirable; dividing by a factor containing the unknown throws away every solution that makes the factor zero. This is worth remembering when solving any equation, quadratic or otherwise.

Problem 2.5.

Solve each equation, taking care not to lose a solution.

  1. x2+5x−6=0x^2 + 5x - 6 = 0;
  2. x(x−3)=6x(x-3) = 6;
  3. x(1−x)=x(2x−1)x(1-x) = x(2x-1).

Completing the Square

Not every quadratic factorises over the rationals, and one that does not can still be rearranged into a square.

Example 2.13.

Solve 2x2−5x+1=02x^2 - 5x + 1 = 0.

Dividing by 22 gives

x2−52x+12=0,that isx2−52x=−12.x^2 - \tfrac52 x + \tfrac12 = 0, \qquad\text{that is}\qquad x^2 - \tfrac52 x = -\tfrac12 .

Half the coefficient of xx is −54-\tfrac54, so adding (54)2=2516\left(\tfrac54\right)^2 = \tfrac{25}{16} to both sides makes the left-hand side a perfect square:

(x−54)2=2516−12=1716.\left(x - \tfrac54\right)^2 = \tfrac{25}{16} - \tfrac12 = \tfrac{17}{16} .

Taking square roots,

x−54=±174,sox=5±174.x - \tfrac54 = \pm\tfrac{\sqrt{17}}{4}, \qquad\text{so}\qquad x = \frac{5 \pm \sqrt{17}}{4} .

Since 17=4.123\sqrt{17} = 4.123 to four significant figures, x=2.28x = 2.28 or x=0.22x = 0.22.

The Quadratic Formula

Completing the square on the general equation gives a formula for the roots once and for all.

Let a,b,ca, b, c be real with a≠0a \neq 0. The solutions of ax2+bx+c=0ax^2 + bx + c = 0 are

x=−b±b2−4ac2a,x = \frac{-b \pm \sqrt{b^2 - 4ac}}{2a},

provided b2−4acb^2 - 4ac is positive or zero. If b2−4acb^2 - 4ac is negative the equation has no solution in the real numbers.

We obtain this by completing the square as in the example. Solving the equation amounts to solving ax2+bx=−cax^2 + bx = -c, and dividing by aa makes this

x2+bax=−ca.x^2 + \frac{b}{a}x = -\frac{c}{a} .

To complete the square on the left we want x2+bax=x2+2sxx^2 + \frac{b}{a}x = x^2 + 2sx, so s=b2as = \dfrac{b}{2a}. Adding s2=b24a2s^2 = \dfrac{b^2}{4a^2} to both sides gives

(x+b2a)2=b24a2−ca=b2−4ac4a2.\left(x + \frac{b}{2a}\right)^2 = \frac{b^2}{4a^2} - \frac{c}{a} = \frac{b^2 - 4ac}{4a^2} .

If b2−4acb^2 - 4ac is negative the right-hand side is negative and cannot be the square of a real number, so the equation has no real solution. If b2−4acb^2 - 4ac is positive or zero we may take the square root, and

x+b2a=±b2−4ac2a,x + \frac{b}{2a} = \pm\frac{\sqrt{b^2-4ac}}{2a},

which rearranges into the formula.

Remark.

The formula is worth committing to memory. Read it aloud like a line of verse: ”xx equals minus bb, plus or minus the square root of bb squared minus four acac, all over two aa.”

The Discriminant

The quantity under the root decides how many roots there are.

Definition 2.14 (Discriminant).

For the equation ax2+bx+c=0ax^2 + bx + c = 0 the number b2−4acb^2 - 4ac is called the discriminant.

If the discriminant is positive, b2−4ac\sqrt{b^2-4ac} can be evaluated and the two signs give two different values: the equation has two real distinct roots. If it is zero, both signs give the same value x=−b2ax = -\tfrac{b}{2a}: the equation has one repeated root, also called equal roots. If it is negative, b2−4ac\sqrt{b^2-4ac} has no real value and the equation has no real roots. To summarise, ax2+bx+c=0ax^2 + bx + c = 0

  1. has two real distinct roots if b2−4ac>0b^2 - 4ac > 0;
  2. has equal roots if b2−4ac=0b^2 - 4ac = 0;
  3. has no real roots if b2−4ac<0b^2 - 4ac < 0.
b² − 4ac > 0b² − 4ac = 0b² − 4ac < 0two rootsone repeated rootno real roots
Figure 2.2. The three cases. The roots are the points where the curve meets the horizontal axis, and the discriminant records how many such points there are.

Example 2.15 (Determining the nature of the roots).

  1. For 4x2−7x+3=04x^2 - 7x + 3 = 0 the discriminant is (−7)2−4(4)(3)=49−48=1>0(-7)^2 - 4(4)(3) = 49 - 48 = 1 > 0, so there are two distinct real roots.
  2. For x2+ax+a2=0x^2 + ax + a^2 = 0 the discriminant is a2−4a2=−3a2a^2 - 4a^2 = -3a^2, which is negative for every a≠0a \neq 0, so there are no real roots; when a=0a = 0 the discriminant is 00 and the equation x2=0x^2 = 0 has the repeated root x=0x = 0.
  3. For x2−px−q2=0x^2 - px - q^2 = 0 the discriminant is (−p)2−4(1)(−q2)=p2+4q2(-p)^2 - 4(1)(-q^2) = p^2 + 4q^2, which is positive unless pp and qq are both zero, so there are two distinct real roots.

Example 2.16.

Find the value of kk for which 2x2−kx+8=02x^2 - kx + 8 = 0 has equal roots.

Equal roots means the discriminant vanishes, so (−k)2−4(2)(8)=0(-k)^2 - 4(2)(8) = 0, that is k2=64k^2 = 64 and k=±8k = \pm 8.

Problem 2.6.

Determine the nature of the roots of the following, without solving them.

  1. x2−6x+9=0x^2 - 6x + 9 = 0;
  2. 3x2+4x+2=03x^2 + 4x + 2 = 0;
  3. x2+kx−1=0x^2 + kx - 1 = 0, for a real constant kk.

Problem 2.7.

Show that the roots of ax2+(a+b)x+b=0ax^2 + (a+b)x + b = 0 are real for all values of aa and bb, and find a relationship between pp and qq for which the roots of px2+qx+1=0px^2 + qx + 1 = 0 are equal.

The Sum and the Product of the Roots

The roots of a quadratic can be described without being found.

Let α\alpha and β\beta be the roots of ax2+bx+c=0ax^2 + bx + c = 0. Then (x−α)(x−β)=0(x-\alpha)(x-\beta) = 0 has the same solutions, and expanding it gives

x2−(α+β)x+αβ=0.x^2 - (\alpha + \beta)x + \alpha\beta = 0 .

Dividing the original equation by aa gives

x2+bax+ca=0.x^2 + \frac{b}{a}x + \frac{c}{a} = 0 .

Both are quadratics in which the coefficient of x2x^2 is 11 and which are satisfied by exactly α\alpha and β\beta, so the remaining coefficients agree:

α+β=−ba,αβ=ca.\alpha + \beta = -\frac{b}{a}, \qquad \alpha\beta = \frac{c}{a} .

An equation may therefore be written as

x2−(sum of roots) x+(product of roots)=0.x^2 - (\text{sum of roots})\,x + (\text{product of roots}) = 0 .

Example 2.17.

The roots of 2x2−3x+6=02x^2 - 3x + 6 = 0 have sum −(−32)=32-\left(-\tfrac32\right) = \tfrac32 and product 62=3\tfrac62 = 3. Conversely, a quadratic whose roots have sum 77 and product 1010 may be written x2−7x+10=0x^2 - 7x + 10 = 0.

Example 2.18 (Building a new equation from an old one).

The roots of 2x2−7x+4=02x^2 - 7x + 4 = 0 are α\alpha and β\beta. Find 1α+1β\dfrac1\alpha + \dfrac1\beta and 1αβ\dfrac{1}{\alpha\beta}, and write down the equation whose roots are 1α\dfrac1\alpha and 1β\dfrac1\beta.

From the equation, α+β=72\alpha + \beta = \tfrac72 and αβ=2\alpha\beta = 2. Expressing the first quantity in terms of these,

1α+1β=α+βαβ=7/22=74,1αβ=12.\frac1\alpha + \frac1\beta = \frac{\alpha+\beta}{\alpha\beta} = \frac{7/2}{2} = \frac74, \qquad \frac{1}{\alpha\beta} = \frac12 .

The required equation has roots summing to 74\tfrac74 and multiplying to 12\tfrac12, so it is

x2−74x+12=0,that is4x2−7x+2=0.x^2 - \tfrac74 x + \tfrac12 = 0, \qquad\text{that is}\qquad 4x^2 - 7x + 2 = 0 .

Remark.

This method works only when each new root depends in the same way on each old one. It applies to roots α2\alpha^2 and β2\beta^2, or 1α\tfrac1\alpha and 1β\tfrac1\beta; it does not apply to roots α+β\alpha + \beta and α−β\alpha - \beta.

Example 2.19.

If α\alpha and β\beta are the roots of x2+3x−2=0x^2 + 3x - 2 = 0, find α3+β3\alpha^3 + \beta^3 and α3β3\alpha^3\beta^3, and write down the equation whose roots are α3\alpha^3 and β3\beta^3.

Here α+β=−3\alpha + \beta = -3 and αβ=−2\alpha\beta = -2. Expanding (α+β)3=α3+3α2β+3αβ2+β3=α3+β3+3αβ(α+β)(\alpha+\beta)^3 = \alpha^3 + 3\alpha^2\beta + 3\alpha\beta^2 + \beta^3 = \alpha^3 + \beta^3 + 3\alpha\beta(\alpha+\beta) and rearranging,

α3+β3=(α+β)3−3αβ(α+β)=(−3)3−3(−2)(−3)=−27−18=−45,\alpha^3 + \beta^3 = (\alpha+\beta)^3 - 3\alpha\beta(\alpha+\beta) = (-3)^3 - 3(-2)(-3) = -27 - 18 = -45,

and α3β3=(αβ)3=(−2)3=−8\alpha^3\beta^3 = (\alpha\beta)^3 = (-2)^3 = -8. So the required equation is x2−(−45)x+(−8)=0x^2 - (-45)x + (-8) = 0, that is

x2+45x−8=0.x^2 + 45x - 8 = 0 .

Example 2.20.

Find the range of values of kk for which x2−2x−k=0x^2 - 2x - k = 0 has real roots. If the roots differ by one, find kk.

Real roots require b2−4ac⩾0b^2 - 4ac \geqslant 0, and (−2)2−4(1)(−k)=4+4k⩾0(-2)^2 - 4(1)(-k) = 4 + 4k \geqslant 0 gives k⩾−1k \geqslant -1.

If the roots differ by one, write them as α\alpha and α+1\alpha + 1. Their sum is 2α+1=−(−2)=22\alpha + 1 = -(-2) = 2, so α=12\alpha = \tfrac12. Their product is α(α+1)=−k\alpha(\alpha+1) = -k, so −k=12⋅32=34-k = \tfrac12 \cdot \tfrac32 = \tfrac34 and k=−34k = -\tfrac34.

Problem 2.8.

The roots of 3x2+2x−1=03x^2 + 2x - 1 = 0 are α\alpha and β\beta. Without solving the equation, write down and simplify the equation whose roots are

  1. 1α\dfrac1\alpha and 1β\dfrac1\beta;
  2. 2α2\alpha and 2β2\beta;
  3. α2\alpha^2 and β2\beta^2.

Logarithms

If we have to find xx in ab=xa^b = x, we can simply work out aba^b. To solve an equation like xa=bx^a = b we need the aa-th root of bb. When the unknown is the index, as in ax=ba^x = b, neither will do, and this is where logarithms come in.

A logarithm is another word for an index, read backwards. Since 23=82^3 = 8, we may say that 33 is the power to which the base 22 must be raised to obtain 88, or that 33 is the logarithm which, with base 22, gives 88. This is written

3=log⁡28.3 = \log_2 8 .

For a positive base, axa^x has its usual real-number meaning for every real xx, agreeing with the rational powers defined in Lesson 1. For a>1a>1 these powers increase as xx increases; for 0<a<10<a<1 they decrease, and in either case they take every positive value exactly once.

Definition 2.21 (Logarithm).

Let a>0a > 0 with a≠1a \neq 1, and let b>0b > 0. The logarithm of bb to base aa, written log⁡ab\log_a b, is the number xx for which

ax=b.a^x = b .

For rational xx, these powers agree with those of Definition 1.29; for other real xx, use the standard extension described above.

The conditions are what make the definition work. When a>0a > 0 and a≠1a \neq 1, the powers axa^x take every positive value, each exactly once, so for b>0b > 0 there is exactly one real number xx with ax=ba^x = b. In particular, two equal powers of aa have equal indices.

Example 2.22.

Since 32=93^2 = 9, the number 22 is the logarithm which with base 33 gives 99, so log⁡39=2\log_3 9 = 2. Since (15)−2=25\left(\tfrac15\right)^{-2} = 25, we have log⁡1/525=−2\log_{1/5} 25 = -2.

The base may be any admissible number. Common logarithms have base 1010, and it is usual to omit the base and write lg⁡\lg for them, so that

lg⁡5=0.6990,lg⁡100=2,lg⁡0.01=−2,\lg 5 = 0.6990, \qquad \lg 100 = 2, \qquad \lg 0.01 = -2,

these saying respectively that 100.6990=510^{0.6990} = 5, that 102=10010^2 = 100 and that 10−2=0.0110^{-2} = 0.01. Any base other than 1010 must be stated.

Since 104=1000010^4 = 10000, lg⁡10000=4\lg 10000 = 4, and a common logarithm gives a rough count of the digits in a number: a whole number with dd digits lies between 10d−110^{d-1} and 10d10^d, so its common logarithm lies between d−1d - 1 and dd. Numerical values of logarithms are found with a calculator, unless the answer can be spotted, as with lg⁡10000\lg 10000.

Two values follow straight from the definition. Since a0=1a^0 = 1, log⁡a1=0\log_a 1 = 0 for every base aa. Since the power to which aa must be raised to give aka^k is kk, log⁡a(ak)=k\log_a\left(a^k\right) = k, and in particular log⁡aa=1\log_a a = 1. Read the other way round, the definition says that

alog⁡ab=b.a^{\log_a b} = b .

The Laws of Logarithms

Three rules cover all the manipulation, and each comes from a rule for indices.

Let b,c>0b,c>0, and let log⁡ab=x\log_a b = x and log⁡ac=y\log_a c = y, so that ax=ba^x = b and ay=ca^y = c.

For the first rule, bc=axay=ax+ybc = a^x a^y = a^{x+y}, so x+y=log⁡abcx + y = \log_a bc, that is

log⁡ab+log⁡ac=log⁡abc.(1)\log_a b + \log_a c = \log_a bc . \tag{1}

For the second, bc=axay=ax−y\dfrac{b}{c} = \dfrac{a^x}{a^y} = a^{x-y}, so x−y=log⁡abcx - y = \log_a \dfrac{b}{c}, that is

log⁡ab−log⁡ac=log⁡abc.(2)\log_a b - \log_a c = \log_a \frac{b}{c} . \tag{2}

For the third, let nn be any real number and put z=log⁡ab nz = \log_a b^{\,n}, so that az=b n=(ax)n=anxa^z = b^{\,n} = (a^x)^n = a^{nx}, whence z=nxz = nx, that is

nlog⁡ab=log⁡ab n.(3)n \log_a b = \log_a b^{\,n} . \tag{3}

The identities (1)(1), (2)(2) and (3)(3) are the three laws of logarithms. Written with a single symbol for the base they read

log⁡xy=log⁡x+log⁡y,log⁡xy=log⁡x−log⁡y,log⁡x y=ylog⁡x.\log xy = \log x + \log y, \qquad \log \frac{x}{y} = \log x - \log y, \qquad \log x^{\,y} = y \log x .

The laws can also be read straight from the identity alog⁡ab=ba^{\log_a b} = b. For the first,

alog⁡ab+log⁡ac=alog⁡ab alog⁡ac=bc=alog⁡abc,a^{\log_a b + \log_a c} = a^{\log_a b}\, a^{\log_a c} = bc = a^{\log_a bc},

and since two equal powers of aa have equal indices, log⁡ab+log⁡ac=log⁡abc\log_a b + \log_a c = \log_a bc. The second follows in the same way, and for the third

a nlog⁡ab=(alog⁡ab)n=b n=alog⁡ab n.a^{\,n \log_a b} = \left(a^{\log_a b}\right)^n = b^{\,n} = a^{\log_a b^{\,n}} .

Remark (A warning).

The base is the same throughout, and there is no law for a sum: log⁡ab+log⁡ac\log_a b + \log_a c is log⁡abc\log_a bc, and it is not log⁡a(b+c)\log_a(b + c). For instance lg⁡2+lg⁡5=lg⁡10=1\lg 2 + \lg 5 = \lg 10 = 1, while lg⁡(2+5)=lg⁡7=0.8451\lg(2 + 5) = \lg 7 = 0.8451. Very few functions ff satisfy f(b+c)=f(b)+f(c)f(b + c) = f(b) + f(c), so the idea that they do is one to get rid of.

Example 2.23.

Using the laws, an expression may be broken into pieces:

log⁡a2b3=log⁡a2−log⁡b3=2log⁡a−3log⁡b.\log \frac{a^2}{b^3} = \log a^2 - \log b^3 = 2\log a - 3\log b .

The base is not specified here; it may be anything, provided it is the same in every term.

Conversely, several logarithms may be gathered into one:

log⁡100−2log⁡50=log⁡100−log⁡502=log⁡1002500=log⁡125.\log 100 - 2\log 50 = \log 100 - \log 50^2 = \log \frac{100}{2500} = \log \frac{1}{25} .

Problem 2.9.

Express each of the following in terms of log⁡a\log a, log⁡b\log b and log⁡c\log c.

  1. log⁡abc\log abc;
  2. log⁡abc\log \dfrac{ab}{c};
  3. log⁡a3bc\log \dfrac{a^3}{\sqrt{bc}}.

Changing the Base

Tables and calculators supply logarithms to base 1010, so a logarithm to another base has to be converted before it can be evaluated. Take log⁡72\log_7 2 and call it xx. Then 7x=27^x = 2, and taking common logarithms of both sides,

xlg⁡7=lg⁡2,sox=lg⁡2lg⁡7=0.30100.8451=0.3562.x \lg 7 = \lg 2, \qquad\text{so}\qquad x = \frac{\lg 2}{\lg 7} = \frac{0.3010}{0.8451} = 0.3562 .

The same argument in general changes base aa to base bb. If log⁡ac=x\log_a c = x then ax=ca^x = c, and taking logarithms to base bb gives xlog⁡ba=log⁡bcx \log_b a = \log_b c, that is

log⁡ac=log⁡bclog⁡ba.\log_a c = \frac{\log_b c}{\log_b a} .

The identity alog⁡ab=ba^{\log_a b} = b gives this in one line as well:

blog⁡bc=c=alog⁡ac=(blog⁡ba)log⁡ac=b log⁡ba⋅log⁡ac,b^{\log_b c} = c = a^{\log_a c} = \left(b^{\log_b a}\right)^{\log_a c} = b^{\,\log_b a \cdot \log_a c},

and equating the indices of bb gives log⁡bc=log⁡ba⋅log⁡ac\log_b c = \log_b a \cdot \log_a c, which rearranges to the same identity.

Taking c=bc = b in the change-of-base identity, and using log⁡bb=1\log_b b = 1, gives the special case

log⁡ab=1log⁡ba.\log_a b = \frac{1}{\log_b a} .

Exponential Equations

An exponential equation is one in which the unknown appears as an index. The third law brings the index down where it can be reached.

Example 2.24.

Solve 5x=105^x = 10.

Taking common logarithms of both sides and using the third law,

xlg⁡5=lg⁡10,sox=lg⁡10lg⁡5=10.6990=1.43x \lg 5 = \lg 10, \qquad\text{so}\qquad x = \frac{\lg 10}{\lg 5} = \frac{1}{0.6990} = 1.43

to three significant figures.

By the definition of a logarithm, the exact solution of ax=ba^x = b is x=log⁡abx = \log_a b, so the solution above is x=log⁡510x = \log_5 10. Taking logarithms is how such a number is evaluated.

Taking logarithms is no help when the unknown appears in a sum of powers, since log⁡(22x+3⋅2x)\log\bigl(2^{2x} + 3 \cdot 2^x\bigr) cannot be simplified. What works instead is a substitution which turns the equation into a quadratic.

Example 2.25.

Solve 22x+3(2x)−4=02^{2x} + 3\left(2^x\right) - 4 = 0.

Since 22x=(2x)22^{2x} = \left(2^x\right)^2, putting y=2xy = 2^x makes the equation

y2+3y−4=0,that is(y+4)(y−1)=0,y^2 + 3y - 4 = 0, \qquad\text{that is}\qquad (y+4)(y-1) = 0,

so y=−4y = -4 or y=1y = 1. There is no real xx with 2x=−42^x = -4, since a power of 22 is positive. From 2x=12^x = 1 we get x=0x = 0, which is the only solution.

Example 2.26.

Solve 9x−12(3x)+27=09^x - 12\left(3^x\right) + 27 = 0.

Putting everything on the same base, 9x=(32)x=(3x)29^x = \left(3^2\right)^x = \left(3^x\right)^2, so the equation is a quadratic in 3x3^x:

(3x)2−12(3x)+27=0,that is(3x−3)(3x−9)=0.\left(3^x\right)^2 - 12\left(3^x\right) + 27 = 0, \qquad\text{that is}\qquad \left(3^x - 3\right)\left(3^x - 9\right) = 0 .

So 3x=33^x = 3 or 3x=93^x = 9, which give x=1x = 1 or x=2x = 2. The last step, going back from the values of 3x3^x to the values of xx, is easily forgotten, and an answer that stops at 3x=33^x = 3 or 99 loses marks in an examination.

Example 2.27.

Solve the simultaneous equations xy=80xy = 80 and lg⁡x−2lg⁡y=1\lg x - 2\lg y = 1.

By the second and third laws the second equation is lg⁡xy2=1\lg \dfrac{x}{y^2} = 1, so xy2=10\dfrac{x}{y^2} = 10 and x=10y2x = 10y^2. Substituting into the first,

10y3=80,soy3=8,y=2,x=40.10y^3 = 80, \qquad\text{so}\qquad y^3 = 8, \quad y = 2, \quad x = 40 .

Example 2.28.

Solve log⁡3x−4log⁡x3+3=0\log_3 x - 4\log_x 3 + 3 = 0.

By the special case of the change-of-base identity, log⁡x3=1log⁡3x\log_x 3 = \dfrac{1}{\log_3 x}, so the equation becomes

log⁡3x−4log⁡3x+3=0.\log_3 x - \frac{4}{\log_3 x} + 3 = 0 .

Putting y=log⁡3xy = \log_3 x and multiplying through by yy gives y2+3y−4=0y^2 + 3y - 4 = 0, that is (y+4)(y−1)=0(y+4)(y-1) = 0, so y=−4y = -4 or y=1y = 1. Hence

x=3−4=181orx=3.x = 3^{-4} = \tfrac{1}{81} \qquad\text{or}\qquad x = 3 .

Problem 2.10.

Solve the following.

  1. 2x=52^x = 5;
  2. 32x−4(3x)+3=03^{2x} - 4\left(3^x\right) + 3 = 0;
  3. the simultaneous equations log⁡2y=2\log_2 y = 2 and xy=8xy = 8.

Polynomials and Division

The polynomials of the first chapter deserve a closer look. A function ff defined for all numbers is a polynomial if there are numbers a0,a1,…,ana_0, a_1, \ldots, a_n such that for all xx

f(x)=anxn+an−1xn−1+⋯+a1x+a0.(1)f(x) = a_n x^n + a_{n-1}x^{n-1} + \cdots + a_1 x + a_0 . \tag{1}

Example 2.29.

The function ff with f(x)=3x5−2x+1f(x) = 3x^5 - 2x + 1 is a polynomial, and f(1)=3−2+1=2f(1) = 3 - 2 + 1 = 2. The function gg with g(x)=12x4+3x2−x+5g(x) = \tfrac12 x^4 + 3x^2 - x + 5 is a polynomial, and

g(2)=12⋅24+3⋅22−2+5=8+12−2+5=23.g(2) = \tfrac12 \cdot 2^4 + 3 \cdot 2^2 - 2 + 5 = 8 + 12 - 2 + 5 = 23 .

Coefficients and Degree

When a polynomial can be written as in (1)(1) we say it is of degree at most nn, and if an≠0a_n \neq 0 we should like to say it has degree nn. Some care is needed before we may. Could the same function also be written

f(x)=bmxm+⋯+b0f(x) = b_m x^m + \cdots + b_0

with a different top power? Could it happen, say, that

7x5−5x4+2x+1=x6−17x3+x+17x^5 - 5x^4 + 2x + 1 = x^6 - 17x^3 + x + 1

for every number xx? Looking at the two sides settles nothing, and with large coefficients there is no easy test. If the answer were yes then the degree could not be defined at all, since in the example we would not know whether to call it 55 or 66. The answer is no, and we come to it below by way of the roots.

Definition 2.30 (Root).

Let ff be a polynomial. A number cc with f(c)=0f(c) = 0 is called a root of ff.

Example 2.31.

Let f(x)=x2−3x+2f(x) = x^2 - 3x + 2. Then f(1)=0f(1) = 0, so 11 is a root, and f(2)=0f(2) = 0, so 22 is a root as well.

Let f(x)=ax2+bx+cf(x) = ax^2 + bx + c with a≠0a \neq 0. If b2−4ac=0b^2 - 4ac = 0 the polynomial has the single root −b2a-\dfrac{b}{2a}, and if b2−4ac>0b^2 - 4ac > 0 it has the two distinct roots

−b+b2−4ac2aand−b−b2−4ac2a.\frac{-b + \sqrt{b^2-4ac}}{2a} \qquad\text{and}\qquad \frac{-b - \sqrt{b^2-4ac}}{2a} .

These are the quadratic equations of the previous chapter, said in the language of roots.

Taking Out a Root

Let ff be a polynomial of degree at most nn and let cc be a root of it. Then there is a polynomial gg of degree at most n−1n-1 with

f(x)=(x−c) g(x)for all x.f(x) = (x-c)\,g(x) \qquad \text{for all } x .

We show this by rewriting ff in powers of x−cx - c rather than powers of xx. Write f(x)=a0+a1x+a2x2+⋯+anxnf(x) = a_0 + a_1x + a_2x^2 + \cdots + a_nx^n and substitute

x=(x−c)+cx = (x-c) + c

for xx throughout. Each kk-th power ((x−c)+c)k\bigl((x-c)+c\bigr)^k expands into a sum of powers of (x−c)(x-c) multiplied by numbers, so there are numbers b0,b1,…,bnb_0, b_1, \ldots, b_n with

f(x)=b0+b1(x−c)+b2(x−c)2+⋯+bn(x−c)nf(x) = b_0 + b_1(x-c) + b_2(x-c)^2 + \cdots + b_n(x-c)^n

for all xx. Setting x=cx = c makes every term after the first vanish, and f(c)=0f(c) = 0, so b0=0b_0 = 0. We may therefore take the factor (x−c)(x-c) out of what is left:

f(x)=(x−c)(b1+b2(x−c)+⋯+bn(x−c)n−1),f(x) = (x-c)\Bigl(b_1 + b_2(x-c) + \cdots + b_n(x-c)^{n-1}\Bigr),

and g(x)=b1+b2(x−c)+⋯+bn(x−c)n−1g(x) = b_1 + b_2(x-c) + \cdots + b_n(x-c)^{n-1} is the polynomial wanted.

Remark (The leading coefficient survives).

Look more closely at that gg. Expanding ((x−c)+c)k\bigl((x-c)+c\bigr)^k produces one term (x−c)k(x-c)^k and others involving lower powers of x−cx-c, so the highest power of x−cx-c to appear anywhere is the nn-th, and it comes only from ((x−c)+c)n\bigl((x-c)+c\bigr)^n, carrying the coefficient ana_n. Hence bn=anb_n = a_n, and gg has the form

g(x)=anxn−1+lower terms.g(x) = a_n x^{n-1} + \text{lower terms} .

How Many Roots There Can Be

Let ff be a polynomial and let a0,…,ana_0, \ldots, a_n be numbers with an≠0a_n \neq 0 such that f(x)=anxn+an−1xn−1+⋯+a0f(x) = a_nx^n + a_{n-1}x^{n-1} + \cdots + a_0 for all xx. Then ff has at most nn roots.

We show this by taking the roots out one at a time. Let c1,c2,…,crc_1, c_2, \ldots, c_r be distinct roots of ff and suppose r⩾nr \geqslant n. Write f(x)=(x−c1)g1(x)f(x) = (x-c_1)g_1(x) with g1g_1 of degree at most n−1n-1. Then

0=f(c2)=(c2−c1) g1(c2),0 = f(c_2) = (c_2 - c_1)\,g_1(c_2),

and c2≠c1c_2 \neq c_1, so g1(c2)=0g_1(c_2) = 0: that is, c2c_2 is a root of g1g_1. The same argument applies to g1g_1, giving g1(x)=(x−c2)g2(x)g_1(x) = (x-c_2)g_2(x) with g2(c3)=0g_2(c_3) = 0 and g2g_2 of degree at most n−2n-2. Continuing in this way until what is left is constant,

f(x)=(x−c1)(x−c2)⋯(x−cn) cf(x) = (x-c_1)(x-c_2)\cdots(x-c_n)\,c

for all xx, and by the remark above c=an≠0c = a_n \neq 0. So if xx is not one of c1,…,cnc_1, \ldots, c_n then no factor vanishes and f(x)≠0f(x) \neq 0. There is therefore no room for a further root.

The Coefficients Are Determined

Suppose a polynomial ff can be written both as

f(x)=anxn+an−1xn−1+⋯+a0andf(x)=bnxn+bn−1xn−1+⋯+b0.f(x) = a_nx^n + a_{n-1}x^{n-1} + \cdots + a_0 \qquad\text{and}\qquad f(x) = b_nx^n + b_{n-1}x^{n-1} + \cdots + b_0 .

Then ai=bia_i = b_i for every ii.

We show this by subtracting. The polynomial

0=f(x)−f(x)=(an−bn)xn+(an−1−bn−1)xn−1+⋯+(a0−b0)0 = f(x) - f(x) = (a_n - b_n)x^n + (a_{n-1}-b_{n-1})x^{n-1} + \cdots + (a_0 - b_0)

takes the value 00 at every number. Put di=ai−bid_i = a_i - b_i and suppose some did_i is non-zero; let mm be the largest index for which this happens, so that

0=dmxm+⋯+d00 = d_mx^m + \cdots + d_0

for all xx, with dm≠0d_m \neq 0. A polynomial written this way has at most mm roots, and this one has every number for a root. That is impossible, so every did_i is zero.

So there is only one way of writing a polynomial in the form (1)(1): the numbers an,…,a0a_n, \ldots, a_0 are determined by the function. They are called the coefficients of ff, the number ana_n is the leading coefficient when an≠0a_n \neq 0, and a0a_0 is the constant term. The worry raised above is settled, and a polynomial with an≠0a_n \neq 0 may safely be said to have degree nn.

Example 2.32.

Let f(x)=4x5−7x3+x−20f(x) = 4x^5 - 7x^3 + x - 20. Its coefficients are

4,0,−7,0,1,−20,4, \quad 0, \quad -7, \quad 0, \quad 1, \quad -20,

where we have written down the coefficient of every power up to the fifth, those of x4x^4 and x2x^2 being zero. The leading coefficient is 44 and the constant term is −20-20, so ff has degree 55.

Remark (Vanishing at a point, and vanishing everywhere).

It often happens that f(x)=0f(x) = 0 for some xx, that is, that ff has a root. This does not make ff the zero polynomial. We call ff the zero polynomial, or say it is identically zero, only when f(x)=0f(x) = 0 for all numbers xx, which happens exactly when all of its coefficients are 00. A polynomial with even one non-zero coefficient is not the zero polynomial, however many roots it may have.

Factorising

The previous chapter determines all the roots of a polynomial of degree 22. For higher degrees it is much harder. Formulas using radicals exist for degrees 33 and 44, and it is a classical result that no such formula can exist in general from degree 55 upwards.

For a polynomial with integer coefficients one may look for the rational roots, and a great deal of time is often spent factoring to find them. It is unusual for a polynomial to factorise so obligingly, and in the quadratic case the systematic answer is the formula rather than a search for factors. Two cases of factorising are worth having.

Example 2.33.

Let f(x)=3x2−5x+1f(x) = 3x^2 - 5x + 1. Its discriminant is 25−12=1325 - 12 = 13, so by the formula its roots are

c1=5+136,c2=5−136,c_1 = \frac{5 + \sqrt{13}}{6}, \qquad c_2 = \frac{5 - \sqrt{13}}{6},

and taking each root out in turn factors ff into two factors of degree 11:

f(x)=3(x−c1)(x−c2).f(x) = 3(x - c_1)(x - c_2) .

The general case runs the same way. If f(x)=ax2+bx+cf(x) = ax^2 + bx + c with a≠0a \neq 0 and b2−4ac>0b^2 - 4ac > 0, then ff has two distinct roots c1c_1 and c2c_2 and

f(x)=a(x−c1)(x−c2).f(x) = a(x-c_1)(x-c_2) .

Example 2.34.

Let f(x)=xn−1f(x) = x^n - 1. Then f(1)=1n−1=0f(1) = 1^n - 1 = 0, so 11 is a root and x−1x - 1 must be a factor:

xn−1=(x−1) g(x)x^n - 1 = (x-1)\,g(x)

for some polynomial gg. Multiplying out (x−1)(xn−1+xn−2+⋯+x+1)(x-1)(x^{n-1} + x^{n-2} + \cdots + x + 1) and watching the terms cancel in pairs identifies gg.

Long Division

Dividing one polynomial by another is the exact analogue of dividing one positive integer by another with a remainder, so we recall that first.

Example 2.35 (Dividing integers).

Divide 327327 by 1717. Set out as at school,

        1 9
    ----------
 17 ) 3 2 7
      1 7
      -----
      1 5 7
      1 5 3
      -----
            4

which tells us that 327=19⋅17+4327 = 19 \cdot 17 + 4, and 44 is the remainder. The first digit 11 was chosen as the largest integer whose product with 1717 is at most 3232; multiplying, subtracting and bringing down the 77 left 157157; then 99 was chosen as the largest integer whose product with 1717 is at most 157157, and subtracting 153153 left 44. Since 4<174 < 17 we stop.

In general, for positive integers nn and dd there are an integer q⩾0q \geqslant 0 and an integer rr with 0⩽r<d0 \leqslant r < d such that n=qd+rn = qd + r. Note that although the procedure is called long division, it uses only multiplication and subtraction.

The same procedure works for polynomials, and is called the Euclidean algorithm: for non-zero polynomials ff and gg there are polynomials qq and rr, the degree of rr being smaller than the degree of gg, such that

f(x)=q(x) g(x)+r(x).f(x) = q(x)\,g(x) + r(x) .

The polynomial rr is the remainder.

Example 2.36 (Dividing polynomials).

Let f(x)=4x3−3x2+x+2f(x) = 4x^3 - 3x^2 + x + 2 and g(x)=x2+1g(x) = x^2 + 1. Laid out as for integers,

              4x -  3
          ------------------------
 x² + 1 ) 4x³ - 3x² +  x + 2
          4x³       + 4x
          ------------------------
              - 3x² - 3x + 2
              - 3x²      - 3
          ------------------------
                    - 3x + 5

so that q(x)=4x−3q(x) = 4x - 3 and r(x)=−3x+5r(x) = -3x + 5, and

4x3−3x2+x+2=(4x−3)(x2+1)+(−3x+5).4x^3 - 3x^2 + x + 2 = (4x-3)(x^2+1) + (-3x+5) .

Each step goes as follows. We take 4x4x first because 4x⋅x24x \cdot x^2 is the term of highest degree in ff; multiplying 4x4x by x2+1x^2 + 1 gives 4x3+4x4x^3 + 4x, which we write underneath with matching powers aligned, and subtracting leaves −3x2−3x+2-3x^2 - 3x + 2. We then take −3-3, because (−3)⋅x2(-3) \cdot x^2 is the term of highest degree in what is left; multiplying gives −3x2−3-3x^2 - 3, and subtracting leaves −3x+5-3x + 5. This has degree 11, smaller than the degree of gg, so the computation is finished.

The reason the procedure works is visible in the steps. Writing f3=ff_3 = f, the multiplier 4x4x was chosen so that f3(x)−4x(x2+1)f_3(x) - 4x(x^2+1) has degree 22, the term 4x34x^3 cancelling; call the result f2f_2. The multiplier −3-3 was then chosen so that f2(x)−(−3)(x2+1)f_2(x) - (-3)(x^2+1) has degree 11, the term −3x2-3x^2 cancelling. Putting the two together,

f3(x)−(4x−3) g(x)=−3x+5,f_3(x) - (4x-3)\,g(x) = -3x + 5,

which is exactly f=qg+rf = qg + r.

Example 2.37.

Let f(x)=2x4−3x2+1f(x) = 2x^4 - 3x^2 + 1 and g(x)=x2−x+3g(x) = x^2 - x + 3. The same pattern gives

                 2x² + 2x -  7
          ---------------------------------
 x² - x + 3 ) 2x⁴       - 3x²      + 1
              2x⁴ - 2x³ + 6x²
          ---------------------------------
                  + 2x³ - 9x²      + 1
                    2x³ - 2x² + 6x
          ---------------------------------
                        - 7x² - 6x + 1
                        - 7x² + 7x - 21
          ---------------------------------
                              - 13x + 22

so q(x)=2x2+2x−7q(x) = 2x^2 + 2x - 7 and r(x)=−13x+22r(x) = -13x + 22.

Remark.

Long division gives a second route to the factor taken out earlier. Dividing ff by x−cx - c,

f(x)=q(x)(x−c)+r(x)f(x) = q(x)(x-c) + r(x)

where rr has degree smaller than 11 and so is a constant, say aa. Evaluating both sides at x=cx = c gives 0=0+a0 = 0 + a, so a=0a = 0: the remainder vanishes and x−cx - c is a factor.

The Remainder on Division

That last observation gives a remainder without any division, so we state it on its own.

If ff is a polynomial which, on division by x−ax - a, gives quotient Q(x)Q(x) and remainder RR, then

f(x)=(x−a) Q(x)+R,f(x) = (x-a)\,Q(x) + R,

and substituting aa for xx makes the first term vanish and gives

R=f(a).R = f(a) .

So when f(x)f(x) is divided by x−ax - a the remainder is f(a)f(a).

Example 2.38.

Divide f(x)=x3−7x2+6x−2f(x) = x^3 - 7x^2 + 6x - 2 by x−2x - 2. Long division gives quotient x2−5x−4x^2 - 5x - 4 and remainder −10-10, that is

x3−7x2+6x−2=(x−2)(x2−5x−4)−10.x^3 - 7x^2 + 6x - 2 = (x-2)(x^2 - 5x - 4) - 10 .

Substituting x=2x = 2 in this identity removes the quotient term and leaves f(2)=−10f(2) = -10, as the rule says.

Example 2.39.

The remainder when x3−2x2+6x^3 - 2x^2 + 6 is divided by x+3x + 3 is

f(−3)=(−3)3−2(−3)2+6=−27−18+6=−39.f(-3) = (-3)^3 - 2(-3)^2 + 6 = -27 - 18 + 6 = -39 .

The remainder when 6x2−7x+26x^2 - 7x + 2 is divided by 2x−12x - 1 is f(12)=64−72+2=0f\left(\tfrac12\right) = \tfrac64 - \tfrac72 + 2 = 0.

Remark.

This gives the remainder only. If the quotient is wanted, the long division must still be carried out.

Recognising a Factor

If x−ax - a is a factor of f(x)f(x) there is no remainder, so R=0R = 0 and f(a)=0f(a) = 0; and conversely, as we saw, a root produces a factor. So

x−a is a factor of f(x)exactly whenf(a)=0.x - a \text{ is a factor of } f(x) \qquad\text{exactly when}\qquad f(a) = 0 .

This lets us factorise polynomials of degree greater than 22 by hand.

Example 2.40 (Factorising a quartic).

Factorise x4−3x3+4x2−8x^4 - 3x^3 + 4x^2 - 8 completely.

Put f(x)=x4−3x3+4x2−8f(x) = x^4 - 3x^3 + 4x^2 - 8. The constant term is −8-8, whose divisors are ±1,±2,±4,±8\pm1, \pm2, \pm4, \pm8, so those are the only whole numbers worth trying. Now

f(1)=1−3+4−8=−6≠0,f(1) = 1 - 3 + 4 - 8 = -6 \neq 0,

so x−1x - 1 is not a factor, while

f(−1)=1+3+4−8=0,f(-1) = 1 + 3 + 4 - 8 = 0,

so x+1x + 1 is a factor. Taking it out, by inspection or by long division,

x4−3x3+4x2−8=(x+1)(x3−4x2+8x−8).x^4 - 3x^3 + 4x^2 - 8 = (x+1)\bigl(x^3 - 4x^2 + 8x - 8\bigr) .

Now put g(x)=x3−4x2+8x−8g(x) = x^3 - 4x^2 + 8x - 8. Since g(−1)=−1−4−8−8≠0g(-1) = -1 - 4 - 8 - 8 \neq 0, the factor x+1x+1 does not repeat; since g(2)=8−16+16−8=0g(2) = 8 - 16 + 16 - 8 = 0, the factor x−2x - 2 appears, and

x3−4x2+8x−8=(x−2)(x2−2x+4).x^3 - 4x^2 + 8x - 8 = (x-2)\bigl(x^2 - 2x + 4\bigr) .

The discriminant of x2−2x+4x^2 - 2x + 4 is 4−16=−12<04 - 16 = -12 < 0, so it has no real roots and no linear factors. Therefore

x4−3x3+4x2−8=(x+1)(x−2)(x2−2x+4).x^4 - 3x^3 + 4x^2 - 8 = (x+1)(x-2)\bigl(x^2 - 2x + 4\bigr) .

Note that after the first factor has been taken out it must be tried again on what remains, since repeated factors are common.

The Factors of a3±b3a^3 \pm b^3

Read as a polynomial in aa, the expression a3−b3a^3 - b^3 vanishes when a=ba = b, so a−ba - b is a factor of it; and a3+b3a^3 + b^3 vanishes when a=−ba = -b, so a+ba + b is a factor of that. Carrying out the division gives

a3−b3=(a−b)(a2+ab+b2),a3+b3=(a+b)(a2−ab+b2),a^3 - b^3 = (a-b)\bigl(a^2 + ab + b^2\bigr), \qquad a^3 + b^3 = (a+b)\bigl(a^2 - ab + b^2\bigr),

as multiplying out confirms.

Problem 2.11.

Determine whether each linear function is a factor of the polynomial beside it.

  1. x−1x - 1 and x3−7x+6x^3 - 7x + 6;
  2. x−2x - 2 and x3−6x2+6x−2x^3 - 6x^2 + 6x - 2;
  3. x+ax + a and x3+ax2−ax−a2x^3 + ax^2 - ax - a^2.

Problem 2.12.

Factorise as far as possible.

  1. x3+2x2−x−2x^3 + 2x^2 - x - 2;
  2. 27x3−127x^3 - 1;
  3. x3+a3x^3 + a^3.

Problem 2.13.

If x2−7x+ax^2 - 7x + a leaves remainder 11 on division by x+1x + 1, find aa. If x−2x - 2 is a factor of ax3−12x+4ax^3 - 12x + 4, find aa.

Binomial Expansions

Definition 2.41 (Binomial).

A binomial is the sum, or difference, of two terms. So a+ba+b, 2x+3y2x+3y and p−rp - r are binomials.

Expanding a power of a binomial by repeated multiplication is quick enough for a square,

(x+y)3=(x+y)(x+y)2=(x2+2xy+y2)(x+y)=x3+3x2y+3xy2+y3,(x+y)^3 = (x+y)(x+y)^2 = (x^2 + 2xy + y^2)(x+y) = x^3 + 3x^2y + 3xy^2 + y^3,

and tedious from the cube upwards. There is a much faster route.

Pascal’s Triangle

Set out the first few powers of 1+x1+x:

(1+x)0=1,(1+x)1=1+x,(1+x)2=1+2x+x2,(1+x)3=1+3x+3x2+x3,(1+x)4=1+4x+6x2+4x3+x4.\begin{aligned} (1+x)^0 &= 1, \\ (1+x)^1 &= 1 + x, \\ (1+x)^2 &= 1 + 2x + x^2, \\ (1+x)^3 &= 1 + 3x + 3x^2 + x^3, \\ (1+x)^4 &= 1 + 4x + 6x^2 + 4x^3 + x^4 . \end{aligned}

Each expansion is the one above it multiplied by (1+x)(1+x), which is to say the one above it added to the one above it shifted up a power. So the coefficient of any power of xx is the sum of the coefficients of the same power and the preceding power in the line before. Writing the coefficients alone as a triangular array, and continuing the rule as far as we please, gives Pascal’s triangle.

111121133114641151010511615201561172135352171
Figure 2.3. Pascal’s triangle. Each entry is the sum of the two above it, as the arrows show for 10+10=2010 + 10 = 20.

Row five of the triangle therefore gives

(1+x)5=1+5x+10x2+10x3+5x4+x5.(1+x)^5 = 1 + 5x + 10x^2 + 10x^3 + 5x^4 + x^5 .

Expanding a Power

Once the coefficients are known, any binomial power follows by substitution.

Example 2.42.

Expand (1+3y)3(1+3y)^3.

From the triangle, (1+x)3=1+3x+3x2+x3(1+x)^3 = 1 + 3x + 3x^2 + x^3. Replacing xx by 3y3y,

(1+3y)3=1+3(3y)+3(3y)2+(3y)3=1+9y+27y2+27y3.(1+3y)^3 = 1 + 3(3y) + 3(3y)^2 + (3y)^3 = 1 + 9y + 27y^2 + 27y^3 .

Example 2.43.

Expand (a+b)4(a+b)^4.

From the triangle, (1+x)4=1+4x+6x2+4x3+x4(1+x)^4 = 1 + 4x + 6x^2 + 4x^3 + x^4. Writing

(a+b)4=a4(1+ba)4(a+b)^4 = a^4\left(1 + \frac{b}{a}\right)^4

and replacing xx by ba\dfrac{b}{a} gives

(a+b)4=a4(1+4ba+6b2a2+4b3a3+b4a4)=a4+4a3b+6a2b2+4ab3+b4.(a+b)^4 = a^4\left(1 + 4\frac{b}{a} + 6\frac{b^2}{a^2} + 4\frac{b^3}{a^3} + \frac{b^4}{a^4}\right) = a^4 + 4a^3b + 6a^2b^2 + 4ab^3 + b^4 .

Three things are worth noticing in that expansion.

  1. The powers of aa and bb in each term add up to 44.
  2. As the powers of aa decrease, the powers of bb increase.
  3. The numerical coefficients are those of (1+x)4(1+x)^4.

With these, an expansion may be written straight down from the triangle. Row six gives

(a+b)6=a6+6a5b+15a4b2+20a3b3+15a2b4+6ab5+b6.(a+b)^6 = a^6 + 6a^5b + 15a^4b^2 + 20a^3b^3 + 15a^2b^4 + 6ab^5 + b^6 .

Example 2.44.

Expand (2x−3y)3(2x-3y)^3.

Using (a+b)3=a3+3a2b+3ab2+b3(a+b)^3 = a^3 + 3a^2b + 3ab^2 + b^3 with a=2xa = 2x and b=−3yb = -3y,

(2x−3y)3=(2x)3+3(2x)2(−3y)+3(2x)(−3y)2+(−3y)3=8x3−36x2y+54xy2−27y3.(2x-3y)^3 = (2x)^3 + 3(2x)^2(-3y) + 3(2x)(-3y)^2 + (-3y)^3 = 8x^3 - 36x^2y + 54xy^2 - 27y^3 .

Problem 2.14.

Expand each of the following.

  1. (1+2x)4(1+2x)^4;
  2. (1−y)5(1-y)^5;
  3. (2x−1)3(2x - 1)^3.

Problem 2.15.

By expanding (1+0.01)3(1+0.01)^3, evaluate (1.01)3(1.01)^3 exactly without a calculator. Then evaluate (2.1)3(2.1)^3 the same way.

Exercises

Exercise 2.1.

Find the stated values.

  1. f(x)=2x+x2−5f(x) = 2x + x^2 - 5: find f(1)f(1) and f(−1)f(-1).
  2. f(x)=x4f(x) = \sqrt[4]{x}: for which numbers does this formula define a function, and what is f(16)f(16)?
  3. f(x)=1x−2f(x) = \dfrac{1}{x-2}: for which numbers does this formula define a function, and what is f(0)f(0)?

Exercise 2.2.

Write ∣x∣=x2|x| = \sqrt{x^2}, the positive square root, so that ∣x∣|x| is xx when xx is positive or zero and −x-x when xx is negative.

  1. Find ∣1∣|1|, ∣−3∣|-3| and ∣−23∣\left|-\tfrac23\right|.
  2. Let g(x)=x+∣x∣g(x) = x + |x|. Find g(2)g(2), g(−4)g(-4) and g(−5)g(-5), and describe in a sentence what gg does to a number.

Exercise 2.3.

A function defined for all numbers is even if f(x)=f(−x)f(x) = f(-x) for every xx, and odd if f(x)=−f(−x)f(x) = -f(-x) for every xx. Determine which of the following are even, which are odd, and which are neither.

  1. f(x)=xf(x) = x;
  2. f(x)=x2f(x) = x^2;
  3. f(x)=x3f(x) = x^3;
  4. f(x)=1/xf(x) = 1/x for x≠0x \neq 0, with f(0)=0f(0) = 0.

Exercise 2.4.

Show that any function ff defined for all numbers can be written as the sum of an even function and an odd function, by considering

f(x)+f(−x)2andf(x)−f(−x)2.\frac{f(x) + f(-x)}{2} \qquad\text{and}\qquad \frac{f(x) - f(-x)}{2} .

Exercise 2.5.

Let f(x)=x2−1f(x) = x^2 - 1 and g(x)=x+1g(x) = x + 1, both defined for all numbers.

  1. Write down formulas for (f+g)(x)(f+g)(x) and (fg)(x)(fg)(x).
  2. Find (f+g)(2)(f + g)(2) and (fg)(−1)(fg)(-1).
  3. Show that ff has a root which gg also has, and factorise ff to display it.

Exercise 2.6.

Express each of the following in partial fractions.

  1. 7(x−3)(x+4)\dfrac{7}{(x-3)(x+4)};
  2. 2x+1(x−1)2\dfrac{2x+1}{(x-1)^2};
  3. x2+x(x+2)(x−1)\dfrac{x^2 + x}{(x+2)(x-1)}.

Exercise 2.7.

Solve each equation, giving surds in exact form.

  1. 3x2−7x+2=03x^2 - 7x + 2 = 0;
  2. x2−4x+1=0x^2 - 4x + 1 = 0;
  3. x(x+2)=3x+6x(x+2) = 3x + 6.

Exercise 2.8.

Determine the nature of the roots of each equation without solving it, and in the last case find the values of kk.

  1. 5x2−2x+1=05x^2 - 2x + 1 = 0;
  2. x2+23 x+3=0x^2 + 2\sqrt{3}\,x + 3 = 0;
  3. kx2+4x+k=0kx^2 + 4x + k = 0 has equal roots.

Exercise 2.9.

The roots of x2−5x+2=0x^2 - 5x + 2 = 0 are α\alpha and β\beta. Without solving the equation, find

  1. α+β\alpha + \beta and αβ\alpha\beta;
  2. α2+β2\alpha^2 + \beta^2;
  3. the equation whose roots are α+1\alpha + 1 and β+1\beta + 1.

Exercise 2.10.

For positive a,b,ca,b,c, with one fixed base throughout, express in terms of log⁡a\log a, log⁡b\log b and log⁡c\log c.

  1. log⁡a2bc\log \dfrac{a^2b}{c};
  2. log⁡ab\log \sqrt{ab};
  3. log⁡1a3\log \dfrac{1}{a^3}.

Exercise 2.11.

Solve the following.

  1. 3x=203^x = 20, to three significant figures;
  2. 4x−5(2x)+4=04^x - 5\left(2^x\right) + 4 = 0;
  3. log⁡2x+log⁡2(x−2)=3\log_2 x + \log_2 (x-2) = 3.

Exercise 2.12.

Show that log⁡ab×log⁡ba=1\log_a b \times \log_b a = 1, and use a change of base to evaluate log⁡512\log_5 12 to three significant figures.

Exercise 2.13.

Let f(x)=2x3−x2−5x−2f(x) = 2x^3 - x^2 - 5x - 2.

  1. Find the remainder when f(x)f(x) is divided by x−1x - 1.
  2. Show that x+1x + 1 is a factor of f(x)f(x), and factorise f(x)f(x) completely.
  3. Write down all the roots of ff.

Exercise 2.14.

Carry out the long division of ff by gg, giving the quotient and the remainder.

  1. f(x)=x3+2x2−x+5f(x) = x^3 + 2x^2 - x + 5 and g(x)=x+3g(x) = x + 3;
  2. f(x)=3x4−x2+2f(x) = 3x^4 - x^2 + 2 and g(x)=x2+x−1g(x) = x^2 + x - 1.

Exercise 2.15.

Factorise as far as possible, using the factors of a3±b3a^3 \pm b^3 where they help.

  1. 8x3−278x^3 - 27;
  2. x3+64x^3 + 64;
  3. x4−16x^4 - 16.

Exercise 2.16.

Expand the following, assuming a≠0a\neq 0 in the third part.

  1. (1+3x)4(1 + 3x)^4;
  2. (x−2y)3(x - 2y)^3;
  3. (a+1a)4\left(a + \tfrac1a\right)^4, simplifying each term.

Exercise 2.17.

Use Pascal’s triangle to write down the expansion of (1+2)4(1+\sqrt{2})^4, and simplify it to the form p+q2p + q\sqrt{2} with pp and qq whole numbers. Do the same for (1−2)4(1-\sqrt{2})^4, and add your two answers.

Check Yourself

 

Fresh questions on the whole chapter — none of them is worked out above. Do each on paper first; the box only tells you whether you got there.

Answers are checked in your browser, as often as you like. Nothing is sent anywhere and nothing is kept but your own work. A formula may be written with the symbols themselves or with ~ & | -> <-> ^, and \and, \or, \to expand as you type.

Exercise 2.18.

Let f(x)=x2−4f(x) = x^2 - 4. What is f(−3)f(-3)?

answer one of these

Exercise 2.19.

What is the degree of the polynomial 7−2x4+x7 - 2x^4 + x?

answer one of these

Exercise 2.20.

Which of these fractional functions is proper?

answer one of these

Exercise 2.21.

In the decomposition 1(x−1)(x+1)=Ax−1+Bx+1\dfrac{1}{(x-1)(x+1)} = \dfrac{A}{x-1} + \dfrac{B}{x+1}, what is AA?

answer one of these

Exercise 2.22.

How many partial fractions does a repeated factor (2x+1)3(2x+1)^3 in a denominator give rise to?

answer one of these

Exercise 2.23.

What is the discriminant of x2−6x+4x^2 - 6x + 4?

answer one of these

Exercise 2.24.

The equation x2+4x+4=0x^2 + 4x + 4 = 0 has

answer one of these

Exercise 2.25.

The roots of 2x2+6x−5=02x^2 + 6x - 5 = 0 are α\alpha and β\beta. What is α+β\alpha + \beta?

answer one of these

Exercise 2.26.

Which equation has roots whose sum is 44 and whose product is −5-5?

answer one of these

Exercise 2.27.

What is log⁡464\log_4 64?

answer one of these

Exercise 2.28.

For positive xx and yy, written as a single logarithm, 2log⁡x−log⁡y2\log x - \log y is

answer one of these

Exercise 2.29.

What is the remainder when x3+2x−1x^3 + 2x - 1 is divided by x−2x - 2?

answer one of these

Exercise 2.30.

For which value of aa is x−3x - 3 a factor of x2−ax+6x^2 - ax + 6?

answer one of these

Exercise 2.31.

At most how many roots can a polynomial of degree 44 have?

answer one of these

Exercise 2.32.

What is the coefficient of x2x^2 in the expansion of (1+x)6(1+x)^6?

answer one of these

Exercise 2.33.

In the expansion of (a+b)5(a+b)^5, what is the coefficient of a2b3a^2b^3?

answer one of these

Lesson 3

Graphs, and Coordinate Geometry

Taught

Functions and Mappings

Mappings

Consider the function ff defined by f:x↦2x+1f : x \mapsto 2x + 1. If we input the value 22 for xx, ff gives 55 as output, and we say that this function maps 22 to 55, written 2↦52 \mapsto 5. Similarly ff maps −2-2 to −3-3, −1-1 to −1-1 and 00 to 11. Here one value of the independent variable xx maps to just one value of the dependent variable f(x)f(x), and no value of f(x)f(x) is reached from two different values of xx.

Now consider f:x↦x2−2x+6f : x \mapsto x^2 - 2x + 6. This function maps −1-1 to 99, 00 to 66, 11 to 55, 22 to 66 and 33 to 99. One value of xx again gives just one value of f(x)f(x), but a value of f(x)f(x) is not necessarily obtained from only one value of xx: both 00 and 22 map to 66.

−2−1012−3−1135x ↦ 2x + 1, one-one−10123569x ↦ x² − 2x + 6, many-one
Figure 3.1. Mapping diagrams for the two functions. On the left every output is reached from exactly one input; on the right 66 and 99 are each reached twice.

Definition 3.1 (One-One and Many-One Mappings).

A mapping is one-one if each output arises from exactly one input, and many-one if at least one output arises from more than one input. A mapping which sends some input to more than one output is one-many.

So x↦2x+1x \mapsto 2x + 1 is a one-one mapping and x↦x2−2x+6x \mapsto x^2 - 2x + 6 a many-one mapping.

We may regard a function as a rule for mapping a number aa to a number bb. The strict meaning of the word, however, is restricted to those relationships in which one input value gives rise to just one output value, and a function in the sense of Definition 2.1 has this built in, since it associates to each number a single other number. Both mappings above are functions. The relationship x↦±xx \mapsto \pm\sqrt{x} is not: input 44 and it gives two outputs, 22 and −2-2. It is one-many, a perfectly good mapping for positive values of xx, and not a function.

Graphs of Functions

Take again f(x)=x2−2x+6f(x) = x^2 - 2x + 6. For any chosen value of xx the corresponding value of f(x)f(x) can be calculated.

xx−1-10011223344
f(x)f(x)99665566991414

Arranged in order of increasing xx, the values of f(x)f(x) show a pattern, and the pattern is easier to see when the results are displayed graphically: we plot the values of f(x)f(x) on a vertical number line against the corresponding values of xx on a horizontal one.

−2−1123451015xy(1, 5)x = 1
Figure 3.2. The curve y=x2−2x+6y = x^2 - 2x + 6, with its values at the integers from −2-2 to 44 marked.

Figure 3.2 was drawn through values of xx one unit apart. Values taken from a larger range, or closer together, all lie on the same smooth curve, so the curve represents every pair of related values of xx and f(x)f(x). From it we can read off properties of the function.

  1. The lowest point on the curve is where x=1x = 1 and f(x)=5f(x) = 5. Although xx, as the independent variable, may be given any real value, the corresponding value of f(x)f(x) is always greater than 55, or equal to 55 when x=1x = 1. We say that the function has a least value of 55.
  2. The curve is symmetrical about the line x=1x = 1: two values of xx equidistant from 11, of the form 1+a1 + a and 1−a1 - a, give the same value of f(x)f(x).
  3. Conversely, every value of f(x)f(x) above the least corresponds to two values of xx, of the form 1+a1 + a and 1−a1 - a.

All three can be confirmed algebraically. Completing the square on the right-hand side gives

f(x)=x2−2x+6=(x−1)2+5.f(x) = x^2 - 2x + 6 = (x - 1)^2 + 5 .

For the first, (x−1)2(x - 1)^2 is a squared quantity and so is never negative. Its least value is 00, which occurs when x=1x = 1, so the least value that f(x)f(x) can have is 55, at x=1x = 1.

For the second, take two values of xx placed symmetrically on either side of x=1x = 1:

f(1+a)=a2+5,f(1−a)=(−a)2+5=a2+5,f(1 + a) = a^2 + 5, \qquad f(1 - a) = (-a)^2 + 5 = a^2 + 5,

so values of xx symmetrical about x=1x = 1 give the same value of f(x)f(x).

For the third, the values of xx giving any particular value cc of f(x)f(x) solve x2−2x+6−c=0x^2 - 2x + 6 - c = 0, and by the quadratic formula the roots of this equation are

x=1+c−5andx=1−c−5.x = 1 + \sqrt{c - 5} \qquad\text{and}\qquad x = 1 - \sqrt{c - 5} .

For these to be real we need c⩾5c \geqslant 5, which is the least value once more, and the two values of xx corresponding to one value of f(x)f(x) are symmetrical about x=1x = 1.

Domain and Range

For f(x)=x2−2x+6f(x) = x^2 - 2x + 6 we have assumed that any real value of xx may be input. A full definition of a function must state the set of permissible input values, so the function illustrated above is fully defined as

f:x↦x2−2x+6,x∈R,f : x \mapsto x^2 - 2x + 6, \quad x \in \RR,

where x∈Rx \in \RR means that xx may be any real number.

Definition 3.2 (Domain and Range).

The set of input values for which a function is defined is its domain. Once the domain is fixed there is a corresponding set of output values, called the range of the function, or its image-set.

The domain is not always the whole of R\RR, as the examples below show. The analysis above also showed that, for the domain x∈Rx \in \RR, the output values were limited: the range of f:x↦x2−2x+6f : x \mapsto x^2 - 2x + 6, x∈Rx \in \RR, is f(x)⩾5f(x) \geqslant 5.

Example 3.3.

Sketch the function f:x↦2x+1f : x \mapsto 2x + 1, x∈Rx \in \RR, and state its range.

The graph is the straight line through (0,1)(0, 1) rising two units for each unit to the right. As xx can be any real number the line is infinite in both directions, so the range of ff is also the whole of R\RR.

Example 3.4.

For the function of the previous example, let AA be the point on the line for which x=2x = 2 and BB the point for which x=−1x = -1. Define the function represented by the segment ABAB, and state its range.

The mapping represented by ABAB is x↦2x+1x \mapsto 2x + 1, but only for values of xx from −1-1 to 22. So the function represented by ABAB is

f:x↦2x+1,x∈R,−1⩽x⩽2,f : x \mapsto 2x + 1, \quad x \in \RR, \quad -1 \leqslant x \leqslant 2,

and its range is the set of values of f(x)f(x) corresponding to −1⩽x⩽2-1 \leqslant x \leqslant 2, that is −1⩽f(x)⩽5-1 \leqslant f(x) \leqslant 5.

Example 3.5.

State the image-set of the function f(x)=2x+1f(x) = 2x + 1, x∈{0,1,2,3,4}x \in \{0, 1, 2, 3, 4\}.

For this function there are just five input values, and the corresponding output values form the image-set, which is {1,3,5,7,9}\{1, 3, 5, 7, 9\}. The graphical representation of this function is not a line but the set of five points.

−2−1123−35xyAB−1 ⩽ x ⩽ 2123413579xyx ∈ {0, 1, 2, 3, 4}
Figure 3.3. One rule on two different domains: the segment ABAB of the second example, and the five points of the third.

Remark.

When the domain of a function is not stated, it is taken to be x∈Rx \in \RR.

Problem 3.1.

  1. Sketch the graph of f:x↦x2f : x \mapsto x^2, x∈Rx \in \RR, and state its range. Then redefine the domain as x⩾0x \geqslant 0 and sketch the graph that now represents ff.
  2. Each of the mappings a↦a2+5a \mapsto a^2 + 5, b↦±bb \mapsto \pm\sqrt{b} and c↦−cc \mapsto -\sqrt{c} has for its domain the non-negative real numbers. State which of them are functions.

The Quadratic Function

Functions of similar form usually have properties in common, so their graphs are usually similar in shape. Knowing the common characteristics of a family of functions lets us sketch any one of them without calculating a table of values.

Definition 3.6 (Quadratic Function).

A function of the form f(x)=ax2+bx+cf(x) = ax^2 + bx + c, where aa, bb and cc are constants and a≠0a \neq 0, is a quadratic function.

The function x2−2x+6x^2 - 2x + 6 analysed above is one. Completing the square on the right-hand side of the general form gives

f(x)=a(x+b2a)2+4ac−b24a.f(x) = a\left(x + \frac{b}{2a}\right)^2 + \frac{4ac - b^2}{4a} .

Whatever value xx takes, 4ac−b24a\dfrac{4ac - b^2}{4a} is constant, equal to KK say, and (x+b2a)2⩾0\left(x + \dfrac{b}{2a}\right)^2 \geqslant 0 as it is a squared quantity. So the function has the general form

f(x)=K+a×(zero or a positive quantity).f(x) = K + a \times (\text{zero or a positive quantity}) .
  1. If aa is positive, f(x)f(x) is at least equal to KK: it has a least value of 4ac−b24a\dfrac{4ac - b^2}{4a}, occurring when x=−b2ax = -\dfrac{b}{2a}.
  2. If aa is negative, f(x)f(x) can never be greater than KK: it has a greatest value of 4ac−b24a\dfrac{4ac - b^2}{4a}, occurring when x=−b2ax = -\dfrac{b}{2a}.

So x=−b2ax = -\dfrac{b}{2a} is the value of xx corresponding to the greatest or least value of f(x)f(x). Taking two values of xx symmetrical about it, x=−b2a±kx = -\dfrac{b}{2a} \pm k,

f(−b2a+k)=ak2+K=f(−b2a−k),f\left(-\frac{b}{2a} + k\right) = ak^2 + K = f\left(-\frac{b}{2a} - k\right),

that is, input values of xx symmetrical about x=−b2ax = -\dfrac{b}{2a} give as output the same value of f(x)f(x). From this analysis we deduce that the curve representing f(x)=ax2+bx+cf(x) = ax^2 + bx + c is symmetrical about the line x=−b2ax = -\dfrac{b}{2a}, called the axis of the curve, and that it turns at a least value when a>0a > 0 and at a greatest value when a<0a < 0.

xyleast valuea > 0xygreatest valuea < 0
Figure 3.4. The two alternative graphs of a quadratic function, each symmetrical about its dashed axis.

The curve representing any particular quadratic function can now be sketched from this information.

Example 3.7 (Sketching a quadratic function).

Sketch the curve representing f(x)=2x2−7x−4f(x) = 2x^2 - 7x - 4.

Here a=2a = 2 and b=−7b = -7, so the axis of the curve is x=−b2a=74x = -\dfrac{b}{2a} = \dfrac{7}{4}, and as a>0a > 0 the function has a least value,

4ac−b24a=4(2)(−4)−(−7)28=−32−498=−818.\frac{4ac - b^2}{4a} = \frac{4(2)(-4) - (-7)^2}{8} = \frac{-32 - 49}{8} = -\frac{81}{8} .

So f(x)f(x) has a least value of −818-\dfrac{81}{8} when x=74x = \dfrac{7}{4}. To locate the curve accurately on the axes we need one more pair of corresponding values, and f(0)f(0) is the easiest to find: f(0)=−4f(0) = -4.

−112345−10−55xy−½4−4(7/4, −81/8)x = 7/4
Figure 3.5. The curve y=2x2−7x−4y = 2x^2 - 7x - 4.

A quicker sketch is available when the function factorises. Here

f(x)=2x2−7x−4=(2x+1)(x−4).f(x) = 2x^2 - 7x - 4 = (2x + 1)(x - 4) .

The coefficient of x2x^2 is positive, so f(x)f(x) has a least value. When f(x)=0f(x) = 0 the corresponding values of xx are the roots of (2x+1)(x−4)=0(2x + 1)(x - 4) = 0, namely x=−12x = -\tfrac12 and x=4x = 4, and the average of these, 74\tfrac74, gives the value of xx about which the curve is symmetrical.

Remark.

This quicker method is suitable only when the function factorises.

Problem 3.2.

Find the greatest or least value of each function, and the value of xx at which it occurs. Then sketch the graph of the third, showing its axis of symmetry clearly.

  1. x2−3x+5x^2 - 3x + 5;
  2. 3−2x−x23 - 2x - x^2;
  3. (1+x)(2−x)(1 + x)(2 - x).

Inequalities

Consider the real numbers 55 and 22, for which 5>25 > 2. The introduction of an extra term on both sides leaves the inequality sign unchanged:

5+4>2+4,that is9>6,and5−4>2−4,that is1>−2.5 + 4 > 2 + 4, \quad\text{that is}\quad 9 > 6, \qquad\text{and}\qquad 5 - 4 > 2 - 4, \quad\text{that is}\quad 1 > -2 .

Multiplication of both sides by a positive number also leaves the sign unchanged: 5×2>2×25 \times 2 > 2 \times 2, that is 10>410 > 4. However, if we multiply both sides by a negative number the inequality is no longer true. Multiplying by −2-2, the left-hand side becomes −10-10 and the right-hand side becomes −4-4, and −10<−4-10 < -4. This illustrates the general fact that multiplication or division of both sides by a negative number reverses the inequality sign.

To summarise, if aa and bb are real numbers such that a>ba > b, then

  1. a+k>b+ka + k > b + k for all real values of kk;
  2. ak>bkak > bk for positive values of kk;
  3. ak<bkak < bk for negative values of kk.

These are the rules for manipulating an inequality.

Example 3.8.

Find the range of values of xx satisfying the inequality x−3<2x+5x - 3 < 2x + 5.

x−3<2x+5x<2x+8adding 3 to both sides−x<8subtracting 2x from both sidesx>−8multiplying by −1\begin{aligned} x - 3 &< 2x + 5 \\ x &< 2x + 8 && \text{adding } 3 \text{ to both sides} \\ -x &< 8 && \text{subtracting } 2x \text{ from both sides} \\ x &> -8 && \text{multiplying by } -1 \end{aligned}

Therefore the range of values of xx that satisfies the inequality is x>−8x > -8.

Problem 3.3.

Find the range of values of xx that satisfies each inequality.

  1. 2x−1<x−42x - 1 < x - 4;
  2. 2(x−1)>3(x−1)2(x - 1) > 3(x - 1);
  3. x+2>4−xx + 2 > 4 - x.

Quadratic Inequalities

Definition 3.9 (Quadratic Inequality).

An inequality that involves a quadratic function is a quadratic inequality.

An example is (x−2)(2x+1)>0(x - 2)(2x + 1) > 0. The range of values of xx satisfying it can be found graphically. Let f(x)=(x−2)(2x+1)f(x) = (x - 2)(2x + 1); the inequality asks for the values of xx at which the curve y=f(x)y = f(x) lies above the xx-axis.

−2−1123−336xy−½2
Figure 3.6. The curve y=(x−2)(2x+1)y = (x - 2)(2x + 1), with the parts above the xx-axis drawn boldly.

From the sketch we see that f(x)>0f(x) > 0, where the curve is above the xx-axis, for values of xx greater than 22 and for values less than −12-\tfrac12. Therefore the ranges of values of xx that satisfy (x−2)(2x+1)>0(x - 2)(2x + 1) > 0 are

x<−12andx>2.x < -\tfrac12 \qquad\text{and}\qquad x > 2 .

Problem 3.4.

Find the range, or ranges, of values of xx that satisfy each inequality.

  1. (x−1)(x−2)>0(x - 1)(x - 2) > 0;
  2. x2−4x>5x^2 - 4x > 5;
  3. (3−2x)(x+5)>0(3 - 2x)(x + 5) > 0.

Problems Involving Quadratic Inequalities

Example 3.10.

Find the range of values of kk for which the equation x2−kx+(k+3)=0x^2 - kx + (k + 3) = 0 has real roots.

Real roots, equal or distinct, require a discriminant which is not negative:

(−k)2−4(k+3)⩾0,that isk2−4k−12⩾0.(-k)^2 - 4(k + 3) \geqslant 0, \qquad\text{that is}\qquad k^2 - 4k - 12 \geqslant 0 .

Let f(k)=k2−4k−12=(k−6)(k+2)f(k) = k^2 - 4k - 12 = (k - 6)(k + 2). Its graph has a least value and meets the axis at k=−2k = -2 and k=6k = 6, and it lies on or above the axis outside these two values. So the equation has real roots for

k⩽−2andk⩾6.k \leqslant -2 \qquad\text{and}\qquad k \geqslant 6 .

Example 3.11.

Find the set of values of pp for which f(x)=x2+3px+pf(x) = x^2 + 3px + p is greater than zero for all real values of xx.

f(x)f(x) is a quadratic function of xx whose coefficient of x2x^2 is positive, so f(x)f(x) has a least value. So if f(x)>0f(x) > 0 for all xx, the least value of f(x)f(x) has to be greater than zero. Completing the square on the right-hand side gives

f(x)=(x+3p2)2+p−9p24,f(x) = \left(x + \frac{3p}{2}\right)^2 + p - \frac{9p^2}{4},

so the least value of f(x)f(x) is p−9p24p - \dfrac{9p^2}{4}, and for f(x)>0f(x) > 0 for all xx we need

p−9p24>0,that is4p−9p2>0,that isp(4−9p)>0.p - \frac{9p^2}{4} > 0, \qquad\text{that is}\qquad 4p - 9p^2 > 0, \qquad\text{that is}\qquad p(4 - 9p) > 0 .

Let g(p)=p(4−9p)g(p) = p(4 - 9p). Its graph has a greatest value and meets the axis at p=0p = 0 and p=49p = \tfrac49, and it is positive between them. Therefore f(x)>0f(x) > 0 for all real xx for the set of values of pp given by 0<p<490 < p < \tfrac49.

Example 3.12.

Find the ranges of values of xx for which 2x−1<x2−4<122x - 1 < x^2 - 4 < 12.

There are two inequality relationships here, and we are looking for the values of xx that satisfy both of them.

  1. 2x−1<x2−42x - 1 < x^2 - 4 gives x2−2x−3>0x^2 - 2x - 3 > 0. Let f(x)=x2−2x−3=(x−3)(x+1)f(x) = x^2 - 2x - 3 = (x - 3)(x + 1); then f(x)>0f(x) > 0 for x<−1x < -1 and x>3x > 3.
  2. x2−4<12x^2 - 4 < 12 gives x2−16<0x^2 - 16 < 0. Let g(x)=x2−16=(x−4)(x+4)g(x) = x^2 - 16 = (x - 4)(x + 4); then g(x)<0g(x) < 0 for −4<x<4-4 < x < 4.

Illustrating these ranges on a number line, the values of xx that satisfy both are where the lines overlap:

−4<x<−1and3<x<4.-4 < x < -1 \qquad\text{and}\qquad 3 < x < 4 .
−5−4−3−2−10123451−5−4−3−2−10123452−5−4−3−2−1012345both
Figure 3.7. The ranges given by the two inequalities, and their overlap.

Some Other Simple Functions

A table of values for each of the following functions gives enough information to deduce the shape of the curve that represents it.

Exponential Functions

Definition 3.13 (Exponential Function).

A function in which the variable appears as an exponent, that is as an index, is an exponential function. The functions 2x2^x, 3x3^x and 10x10^x are all exponential functions of xx.

Consider the function f(x)=2xf(x) = 2^x, for which the following table shows corresponding values of xx and f(x)f(x).

xx−10-10−5-5−4-4−3-3−2-2−1-10011223344551010
2x2^x11024\tfrac{1}{1024}132\tfrac{1}{32}116\tfrac{1}{16}18\tfrac1814\tfrac1412\tfrac12112244881616323210241024

From this table we see that

  1. f(x)>0f(x) > 0 for all real values of xx;
  2. as xx increases, f(x)f(x) increases at a rapidly accelerating rate;
  3. f(x)=1f(x) = 1 when x=0x = 0;
  4. as xx decreases, through x=−10,−100,…x = -10, -100, \ldots, f(x)f(x) quickly becomes numerically smaller, and we say that as xx approaches minus infinity, f(x)f(x) approaches the value zero. This is written x→−∞x \to -\infty, f(x)→0f(x) \to 0.

From these observations the sketch of f(x)=2xf(x) = 2^x is drawn.

−4−3−2−11231248xyasymptote
Figure 3.8. The curve y=2xy = 2^x and its asymptote, the xx-axis.

The curve approaches the xx-axis but never actually touches it or crosses it.

Definition 3.14 (Asymptote).

A line which a curve approaches more and more closely, without ever reaching it, is an asymptote to the curve.

So the xx-axis is an asymptote to the curve y=2xy = 2^x. Another way of expressing the behaviour of f(x)f(x) for negative values of xx is that f(x)f(x) approaches a limiting value, or limit, of zero as xx approaches minus infinity. This property is written

lim⁡x→−∞2x=0.\lim_{x \to -\infty} 2^x = 0 .

Any function of the form axa^x, where a>1a > 1, is represented by a curve similar to that deduced for 2x2^x.

Rational Functions

Definition 3.15 (Rational Function).

A function whose numerator and denominator are both polynomials is a rational function. These are the fractional functions of the last lesson, under their other name.

For example 1x\dfrac{1}{x}, xx−2\dfrac{x}{x - 2} and x2−7x+1\dfrac{x^2 - 7}{x + 1} are rational functions of xx. Consider the function f(x)=1xf(x) = \dfrac{1}{x} and the following table of corresponding values.

xx−3-3−2-2−1-100112233441010
1x\tfrac{1}{x}−13-\tfrac13−12-\tfrac12−1-1undefined1112\tfrac1213\tfrac1314\tfrac14110\tfrac{1}{10}

From this table we see that

  1. for x>0x > 0, f(x)>0f(x) > 0, and as x→∞x \to \infty, f(x)→0f(x) \to 0;
  2. for x<0x < 0, f(x)<0f(x) < 0, and as x→−∞x \to -\infty, f(x)→0f(x) \to 0, so that xx and f(x)f(x) always have the same sign;
  3. for x=0x = 0, f(0)=10f(0) = \dfrac10, which has no finite value and is said to be undefined.

In such circumstances we investigate the behaviour of f(x)f(x) as xx approaches zero. Now xx can approach zero in two ways: it can decrease from positive values towards zero, which is to approach zero from above, or increase from negative values towards zero, which is to approach zero from below.

xx110.10.10.010.010.0010.001
1x\tfrac{1}{x}11101010010010001000
xx−1-1−0.1-0.1−0.01-0.01−0.001-0.001
1x\tfrac{1}{x}−1-1−10-10−100-100−1000-1000

From these tables we see that as xx decreases to zero, f(x)→∞f(x) \to \infty, and as xx increases to zero, f(x)→−∞f(x) \to -\infty. From these observations we can draw the sketch representing f(x)=1xf(x) = \dfrac{1}{x}.

−4−3−2−11234−4−224xy
Figure 3.9. The curve y=1xy = \dfrac{1}{x}. Both axes are asymptotes.

This curve has two asymptotes, the horizontal and the vertical axes. All the curves looked at so far have been unbroken, or continuous, but this curve has a break, or discontinuity, at the point where x=0x = 0 and f(0)f(0) is undefined. In general, if f(x)f(x) is undefined for a finite value of xx, x=ax = a say, then the curve representing f(x)f(x) has a discontinuity where x=ax = a.

The function 1x\dfrac{1}{x} does not approach a unique value as xx approaches zero: it goes to ∞\infty or to −∞-\infty depending on whether xx approaches zero from above or from below. In this case we say that lim⁡x→01x\displaystyle\lim_{x \to 0} \frac{1}{x} does not exist. In general, for lim⁡x→af(x)\displaystyle\lim_{x \to a} f(x) to exist with a value kk, say, f(x)f(x) must approach kk both as xx approaches aa from above and as xx approaches aa from below.

Logarithmic Functions

For values of a>0a > 0 with a≠1a \neq 1, any function of the form log⁡ax\log_a x, log⁡a(2x+1)\log_a (2x + 1), and so on, is a logarithmic function, the logarithm being that of Definition 2.21. Consider the function f(x)=log⁡2xf(x) = \log_2 x and the following table of corresponding values.

xx−1-10014\tfrac1412\tfrac1211224488
log⁡2x\log_2 xdoes not existundefined−2-2−1-100112233

From this table we see that

  1. when x=−1x = -1, f(−1)=log⁡2(−1)=bf(-1) = \log_2(-1) = b say; but there is no real value of bb for which 2b=−12^b = -1, since every power of 22 is positive. So f(−1)f(-1) does not exist, and all negative values of xx lead to the same conclusion: log⁡2x\log_2 x does not exist for negative values of xx;
  2. for x>1x > 1, f(x)>0f(x) > 0, and as x→∞x \to \infty, f(x)→∞f(x) \to \infty;
  3. for x=0x = 0, f(0)f(0) is undefined, so we investigate the behaviour of f(x)f(x) as xx approaches zero from above. As xx decreases to zero, f(x)→−∞f(x) \to -\infty, and for 0<x<10 < x < 1, f(x)<0f(x) < 0.

From these observations we can sketch the graphical representation of f(x)=log⁡2xf(x) = \log_2 x.

1248−3−2−1123xy
Figure 3.10. The curve y=log⁡2xy = \log_2 x.

Any function of the form log⁡ax\log_a x with a>1a > 1 has a graph of similar shape. Note that, while log⁡ax\log_a x does not exist for negative values of xx, the value of log⁡ax\log_a x can itself be negative.

Simple Variations of Functions

The curve representing f(x)=2xf(x) = 2^x is now familiar, and from it we can obtain the graphs of some simple variations.

  1. g(x)=2x+1g(x) = 2^x + 1. Comparing g(x)=2x+1g(x) = 2^x + 1 with f(x)=2xf(x) = 2^x, for any one value of xx, g(x)g(x) is one unit greater than f(x)f(x). So the curve representing g(x)g(x) is the same shape as that representing f(x)f(x) but raised vertically by one unit. In general, the curve representing f(x)+cf(x) + c is the curve representing f(x)f(x) raised vertically by cc units.
  2. g(x)=2−xg(x) = 2^{-x}. For g(x)g(x) an input x=ax = a gives the output 2−a2^{-a}, while for f(x)f(x) the input x=−ax = -a gives the same output. So g(a)=f(−a)g(a) = f(-a), that is g(x)=f(−x)g(x) = f(-x) for x∈Rx \in \RR, and the curve representing g(x)g(x) is the same as that representing f(x)f(x) with the negative and positive values of xx transposed. In general, the curve representing f(−x)f(-x) is the reflection in the vertical axis of the curve representing f(x)f(x).
  3. g(x)=−2xg(x) = -2^x. Here g(x)=−f(x)g(x) = -f(x), so the curve representing g(x)g(x) is the same shape as that for f(x)f(x) with the positive and negative values of f(x)f(x) transposed. In general, the curve representing −f(x)-f(x) is the reflection in the horizontal axis of the curve representing f(x)f(x).
xyy = 2x + 1xyy = 2−xxyy = −2x
Figure 3.11. The three variations of y=2xy = 2^x, each drawn against the faint original.

Problem 3.5.

Write down the values of f(x)=(12)xf(x) = \left(\tfrac12\right)^x corresponding to x=2,4,6x = 2, 4, 6 and to x=−2,−4,−6x = -2, -4, -6. From these values deduce the behaviour of f(x)f(x) as x→∞x \to \infty and as x→−∞x \to -\infty, and sketch the graph of f(x)f(x), marking any asymptote. Which variation of 2x2^x is this curve?

Inverse Functions

Consider the mapping f:x↦2xf : x \mapsto 2x, x∈{2,3,4}x \in \{2, 3, 4\}. Under this function the domain {2,3,4}\{2, 3, 4\} maps to the image-set {4,6,8}\{4, 6, 8\}.

It is possible to reverse this mapping: we can map each member of the image-set {4,6,8}\{4, 6, 8\} back to the corresponding member of the domain by halving it, 4↦24 \mapsto 2, 6↦36 \mapsto 3, 8↦48 \mapsto 4. Expressed as an algebraic relationship, if x∈{4,6,8}x \in \{4, 6, 8\} then x↦12xx \mapsto \tfrac12 x maps 4↦24 \mapsto 2, 6↦36 \mapsto 3, 8↦48 \mapsto 4.

This reverse mapping is a one-one mapping, so it is a function in its own right, and it is called the inverse function of ff. Denoting this inverse function by f−1f^{-1}, we see that f−1:x↦12xf^{-1} : x \mapsto \tfrac12 x, x∈{4,6,8}x \in \{4, 6, 8\}, reverses the mapping f:x↦2xf : x \mapsto 2x, x∈{2,3,4}x \in \{2, 3, 4\}. In fact f:x↦2xf : x \mapsto 2x can be reversed for all real values of xx, so if ff is the function defined by f:x↦2xf : x \mapsto 2x, x∈Rx \in \RR, then f−1f^{-1} is the function which reverses this mapping, defined by

f−1:x↦12x,x∈R.f^{-1} : x \mapsto \tfrac12 x, \quad x \in \RR .

Now consider the function f:x↦x2f : x \mapsto x^2, x∈Rx \in \RR. This is a many-one mapping: there are two values of xx which map to one value of f(x)f(x), both 22 and −2-2 mapping to 44, for example. We can reverse this mapping by taking the positive and the negative square root of each member of the image-set, which in algebraic form is the mapping x↦±xx \mapsto \pm\sqrt{x}. However, this is a one-many mapping, so it is not a function, and we say that f:x↦x2f : x \mapsto x^2, x∈Rx \in \RR, does not have an inverse function.

If, however, we restrict the domain of ff to x⩾0x \geqslant 0, redefining the function as f:x↦x2f : x \mapsto x^2, x⩾0x \geqslant 0, then it becomes a one-one mapping. The reverse mapping x↦xx \mapsto \sqrt{x} is also one-one, so the function f:x↦x2f : x \mapsto x^2, x⩾0x \geqslant 0, does have an inverse, namely

f−1:x↦x,x⩾0.f^{-1} : x \mapsto \sqrt{x}, \quad x \geqslant 0 .

Definition 3.16 (Inverse Function).

A function ff maps the domain of ff to the image-set of ff. If the reverse mapping, of the image-set of ff to the domain of ff, is a function, it is called the inverse function of ff and is denoted by f−1f^{-1}.

Remark.

If ff defines a one-one mapping then f−1f^{-1} exists, but if ff defines a many-one mapping then f−1f^{-1} does not exist.

The Graphs of Functions and Their Inverses

The graphs of f(x)=2xf(x) = 2x and f−1(x)=12xf^{-1}(x) = \tfrac12 x for x∈Rx \in \RR, and of f(x)=x2f(x) = x^2 and f−1(x)=xf^{-1}(x) = \sqrt{x} for x⩾0x \geqslant 0, are sketched below. Observing these, we see that in each case the graph of f−1(x)f^{-1}(x) is the reflection of the graph of f(x)f(x) in the line y=xy = x.

24682468xyy = 2x and y = ½x12341234xyy = x², x ⩾ 0, and y = √x
Figure 3.12. Two functions and their inverses, each pair reflected in the dashed line y=xy = x.

This is true for the graph of any function ff and its inverse f−1f^{-1}. The line y=xy = x is the graph of f:x↦xf : x \mapsto x, and this function is its own inverse. So if the graph representing a function is known, the graph representing its inverse can be sketched. Even when a function ff does not possess an inverse, the curve representing ff can still be reflected in the line y=xy = x, but in that case the reflected graph does not represent a function.

Example 3.17.

Given the function f(x)=2xf(x) = 2^x, x∈Rx \in \RR, find f−1f^{-1} as a function of xx and sketch the graph of f−1f^{-1}.

The function ff maps xx to 2x2^x. To find f−1f^{-1} we have to reverse this process, that is map values of 2x2^x back to values of xx. If 2x=w2^x = w, say, then taking logarithms to base 22 of each side gives x=log⁡2wx = \log_2 w. Hence the relationship x↦wx \mapsto w can be expressed as w↦log⁡2ww \mapsto \log_2 w, and this is a one-one mapping for w>0w > 0, so it is a function, and it reverses the mapping x↦2xx \mapsto 2^x. Replacing the variable ww by the variable xx, the inverse of f(x)=2xf(x) = 2^x is

f−1(x)=log⁡2x,x>0.f^{-1}(x) = \log_2 x, \quad x > 0 .
−224−224xyy = 2xy = log₂ xy = x
Figure 3.13. The curve y=2xy = 2^x and its inverse y=log⁡2xy = \log_2 x, reflected in the dashed line y=xy = x.

In the same way, for any base a>0a > 0 with a≠1a \neq 1, the function log⁡ax\log_a x from the positive numbers to the real numbers is the inverse of the function axa^x from the real numbers to the positive numbers.

Problem 3.6.

Each of the following functions has the domain x∈Rx \in \RR. Determine which of them have an inverse, and where f−1f^{-1} exists express it as a function of xx.

  1. 3x3x;
  2. 2x+12x + 1;
  3. x2−4x^2 - 4.

Coordinate Geometry

Locating a Point

Graphical methods lend themselves particularly well to the investigation of the geometric properties of many kinds of curves. We restrict ourselves to plane figures, those that can be described fully using only two dimensions, and to represent any figure on a graph we need, as a start, a simple and unambiguous way of describing the position of a point.

Consider the problem of describing the location of a town, Birmingham say. There are many ways in which this can be done, but all require reference to at least one known place and known directions, called a system, or frame, of reference. Within this frame of reference two measurements, or coordinates, are needed to locate the town precisely.

NOB50 km36°52′(i) distance and bearingNOBE30 km40 km(ii) east and north
Figure 3.14. The position of BB described in two alternative ways.

In the first description the system of reference is the fixed point OO and the direction due north from OO, and the coordinates of BB are 5050 km from OO on a bearing of 36∘52′36^\circ 52'. In the second the system of reference is the pair of directions due east and due north from the fixed point OO, and the coordinates of BB are 3030 km east of OO and 4040 km north of OO. The two systems most often used for mathematical analysis are basically similar to these two practical systems.

Polar Coordinates

The system of reference is a fixed point OO, called the pole, and a fixed direction from OO, the line OxOx, called the initial line. The coordinates of a point PP are the distance of PP from OO and the angle OPOP makes with OxOx, measured in an anticlockwise sense from OxOx. These coordinates are written as an ordered pair (r,θ)(r, \theta), that is (distance, angle).

Cartesian Coordinates

The system of reference is a fixed point OO, the origin, and a pair of perpendicular lines through OO. It is usual to draw these lines horizontally and vertically; the horizontal line is called the xx-axis and the vertical line the yy-axis. The coordinates of a point PP are the directed distances of PP from OO parallel to the axes: a positive coordinate is a distance measured in the positive direction of the axis, and a negative coordinate is a distance in the opposite direction. The coordinates are given as an ordered pair (a,b)(a, b), with the xx-coordinate, or abscissa, first and the yy-coordinate, or ordinate, second.

xO(2, 30°)30°2polar−2−112−2−11234xy(1, 4)(−2, −1)Cartesian
Figure 3.15. The point (2,30∘)(2, 30^\circ) in polar coordinates, and the points (1,4)(1, 4) and (−2,−1)(-2, -1) in Cartesian coordinates.

Coordinate geometry is the name given to the analysis, using graphical methods, of geometric properties. The properties of straight lines and of many curves are most simply found using Cartesian coordinates, so this system of reference is used more frequently than any other. For this analysis we need to refer to three types of points:

  1. fixed points whose coordinates are known, such as the point (4,5)(4, 5);
  2. fixed points whose coordinates are not known numerically, referred to as the points (x1,y1)(x_1, y_1), (x2,y2)(x_2, y_2), and so on, or (a,b)(a, b);
  3. points which are not fixed, called general points, a general point being referred to as the point (x,y)(x, y).

It is conventional to use the letters PP, QQ, RR for general points and AA, BB, CC for fixed points. To avoid distorting the shape of a curve when drawing it on a Cartesian plane, the two axes are graduated using identical scales.

Problem 3.7.

Represent on one diagram the points whose polar coordinates are (1,45∘)(1, 45^\circ), (3,90∘)(3, 90^\circ), (1,150∘)(1, 150^\circ) and (2,200∘)(2, 200^\circ), and on another the points whose Cartesian coordinates are (4,2)(4, 2), (−1,5)(-1, 5), (0,3)(0, 3) and (−2,−5)(-2, -5).

The Length of the Line Joining Two Points

12341234xyA(1, 2)B(3, 4)N
Figure 3.16. The points A(1,2)A(1, 2) and B(3,4)B(3, 4), with N(3,2)N(3, 2) completing a right-angled triangle.

From the diagram we see that the length of the line joining A(1,2)A(1, 2) and B(3,4)B(3, 4) can be found by Pythagoras, using the point N(3,2)N(3, 2):

AB2=AN2+BN2=(3−1)2+(4−2)2=8,soAB=8=22.AB^2 = AN^2 + BN^2 = (3 - 1)^2 + (4 - 2)^2 = 8, \qquad\text{so}\qquad AB = \sqrt{8} = 2\sqrt{2} .

In general, if A(x1,y1)A(x_1, y_1) and B(x2,y2)B(x_2, y_2) are any two points, the point N(x2,y1)N(x_2, y_1) gives AN=x2−x1AN = x_2 - x_1 and NB=y2−y1NB = y_2 - y_1, and by Pythagoras AB2=AN2+NB2AB^2 = AN^2 + NB^2. Therefore the length of the line joining A(x1,y1)A(x_1, y_1) to B(x2,y2)B(x_2, y_2) is

AB=(x2−x1)2+(y2−y1)2.AB = \sqrt{(x_2 - x_1)^2 + (y_2 - y_1)^2} .

This formula still holds when some, or all, of the coordinates are negative. For A(−2,2)A(-2, 2) and B(3,−1)B(3, -1), AN=3−(−2)=5AN = 3 - (-2) = 5 and NB=−1−2=−3NB = -1 - 2 = -3, and AB=25+9=34AB = \sqrt{25 + 9} = \sqrt{34}.

The Midpoint of the Line Joining Two Points

1234123456xyA(1, 1)B(3, 5)M(2, 3)
Figure 3.17. The midpoint M(2,3)M(2, 3) of the line joining A(1,1)A(1, 1) and B(3,5)B(3, 5).

Let MM be the midpoint of the line joining A(1,1)A(1, 1) and B(3,5)B(3, 5). The foot of the perpendicular from MM to the xx-axis lies halfway between the feet of the perpendiculars from AA and BB, so the xx-coordinate of MM is

1+12(3−1)=12(3+1)=2.1 + \tfrac12(3 - 1) = \tfrac12(3 + 1) = 2 .

Similarly the yy-coordinate of MM is 1+12(5−1)=12(5+1)=31 + \tfrac12(5 - 1) = \tfrac12(5 + 1) = 3. Therefore MM is the point (2,3)(2, 3).

In general, if MM is the midpoint of the line joining A(x1,y1)A(x_1, y_1) and B(x2,y2)B(x_2, y_2), the xx-coordinate of MM is x1+12(x2−x1)=12(x1+x2)x_1 + \tfrac12(x_2 - x_1) = \tfrac12(x_1 + x_2), the arithmetic mean of the xx-coordinates of AA and BB, and likewise its yy-coordinate is the arithmetic mean of the yy-coordinates. So the coordinates of MM are

(x1+x22,  y1+y22).\left( \frac{x_1 + x_2}{2}, \; \frac{y_1 + y_2}{2} \right) .

This formula also holds when some, or all, of the coordinates are negative: the midpoint of the line joining A(−3,−2)A(-3, -2) and B(1,3)B(1, 3) is (12(−3+1),12(−2+3))=(−1,12)\left(\tfrac12(-3 + 1), \tfrac12(-2 + 3)\right) = \left(-1, \tfrac12\right).

Problem 3.8.

  1. Find the length of the line joining (1,2)(1, 2) and (4,6)(4, 6), and of the line joining (−1,−4)(-1, -4) and (−3,−2)(-3, -2).
  2. Find the coordinates of the midpoints of the same two lines.
  3. MM is the midpoint of ABAB. If AA is (5,7)(5, 7) and MM is (0,2)(0, 2), find the coordinates of BB.

Gradient

Definition 3.18 (Gradient).

The gradient of a straight line is a measure of its slope with respect to the xx-axis. It is the increase in the yy-coordinate divided by the increase in the xx-coordinate between one point on the line and another point on the line.

Consider the line passing through the points A(2,3)A(2, 3) and B(6,1)B(6, 1). From AA to BB the yy-coordinate decreases by 22, that is increases by −2-2, and the xx-coordinate increases by 44, so the gradient of ABAB is −24=−12\dfrac{-2}{4} = -\dfrac12. From BB to AA the gradient is

increase in yincrease in x=2−4=−12,\frac{\text{increase in } y}{\text{increase in } x} = \frac{2}{-4} = -\frac12,

so it does not matter in which order the two points are considered, provided they are considered in the same order when calculating the increases in both xx and yy.

Consider now the straight line through A(1,2)A(1, 2) and B(4,3)B(4, 3). From AA to BB the increase in the yy-coordinate is 11 and the increase in the xx-coordinate is 33, so the gradient of ABAB is 13\dfrac13. If CC and DD are any two other points on the same line, the right-angled triangle they make with lines parallel to the axes is similar to the one made by AA and BB, so it gives the same ratio: the gradient of a line may be found from any two points on the line.

From the two examples we see that the gradient of a line may be positive or negative. A positive gradient indicates an uphill slope with respect to the positive direction of the xx-axis, that is a line which makes an acute angle with the positive sense of the xx-axis. A negative gradient indicates a downhill slope, a line which makes an obtuse angle with it.

xyriserunpositive gradientxyfallrunnegative gradient
Figure 3.18. A positive gradient rises to the right; a negative gradient falls.

In general, the gradient of the line passing through A(x1,y1)A(x_1, y_1) and B(x2,y2)B(x_2, y_2) is

the increase in the y-coordinatethe increase in the x-coordinate=y2−y1x2−x1.\frac{\text{the increase in the } y\text{-coordinate}}{\text{the increase in the } x\text{-coordinate}} = \frac{y_2 - y_1}{x_2 - x_1} .

As the gradient of a straight line is the increase in yy divided by the increase in xx from one point on the line to another, the gradient measures the increase in yy per unit increase in xx, that is the rate of increase of yy with respect to xx.

Parallel Lines

If l1l_1 and l2l_2 are parallel lines, they are equally inclined to the positive direction of the xx-axis, so the gradient triangles of the two lines are similar and give the same ratio: parallel lines have equal gradients.

Perpendicular Lines

Consider two perpendicular lines whose gradients are m1m_1 and m2m_2. Draw through the origin a line OSOS parallel to the first and a line OROR parallel to the second, drop SS to TT on the xx-axis and RR to QQ on the yy-axis. Because OSOS and OROR are perpendicular, the triangle OQROQR is a copy of the triangle OTSOTS turned through a right angle, so the triangles are similar and

STOT=QROQ.\frac{ST}{OT} = \frac{QR}{OQ} .
xySTRQO
Figure 3.19. Perpendicular lines OSOS and OROR, and the similar triangles OTSOTS and OQROQR.

But the gradient of OSOS is STOT=m1\dfrac{ST}{OT} = m_1, and as OROR rises to the left the gradient of OROR is −OQQR=m2-\dfrac{OQ}{QR} = m_2. Since STOT\dfrac{ST}{OT} and OQQR\dfrac{OQ}{QR} are reciprocals,

m1m2=STOT×(−OQQR)=−1.m_1 m_2 = \frac{ST}{OT} \times \left(-\frac{OQ}{QR}\right) = -1 .

So the product of the gradients of perpendicular lines is −1-1, or, if one line has a gradient mm, the gradient of any line perpendicular to it is −1m-\dfrac{1}{m}.

Example 3.19 (A point on a median).

Show that the point (−67,0)\left(-\tfrac67, 0\right) is on the median through AA of triangle ABCABC, where AA, BB, CC are the points (2,4)(2, 4), (−2,3)(-2, 3), (1,−2)(1, -2). If also the point (a,b)(a, b) is on this median, find a relationship between aa and bb.

If ADAD is the median through AA, then DD is the midpoint of BCBC, that is the point (−12,12)\left(-\tfrac12, \tfrac12\right). If E(−67,0)E\left(-\tfrac67, 0\right) is on ADAD, the gradients of ADAD and AEAE should be equal. Now

gradient of AD=4−122+12=75,gradient of AE=4−02+67=75,\text{gradient of } AD = \frac{4 - \tfrac12}{2 + \tfrac12} = \frac75, \qquad \text{gradient of } AE = \frac{4 - 0}{2 + \tfrac67} = \frac75,

so EE is on the median ADAD. A condition that P(a,b)P(a, b) should be on ADAD is that the gradient of APAP equals the gradient of ADAD:

4−b2−a=75,so20−5b=14−7a,that is7a−5b+6=0.\frac{4 - b}{2 - a} = \frac75, \qquad\text{so}\qquad 20 - 5b = 14 - 7a, \qquad\text{that is}\qquad 7a - 5b + 6 = 0 .

Problem 3.9.

  1. Find the gradients of the lines passing through (−1,−3)(-1, -3) and (−2,1)(-2, 1), and through (3,−2)(3, -2) and (−1,4)(-1, 4).
  2. Determine, by comparing gradients, whether the points (0,−1)(0, -1), (1,1)(1, 1) and (2,3)(2, 3) are collinear, that is whether they lie on the same straight line.
  3. Determine whether ABAB is parallel or perpendicular to CDCD, where AA, BB, CC, DD are (0,−1)(0, -1), (1,1)(1, 1), (1,5)(1, 5), (−1,1)(-1, 1).

Equations and Regions

The Cartesian system of reference provides a means of defining the position of any point in a plane, and this plane is called the xyxy plane. In general xx and yy are independent variables, each able to take any value independently of the value of the other, unless some restriction is placed on them.

Consider the set of points for which x=2x = 2. As the value of yy is not restricted, these points all lie on the line parallel to the yy-axis passing through (2,0)(2, 0). So the equation x=2x = 2 defines this line in the xyxy plane; x=2x = 2 is called the equation of the line, which is briefly referred to as “the line x=2x = 2”.

Now consider the set of points for which x>2x > 2. All points to the right of the line x=2x = 2 have an xx-coordinate greater than 22, so the inequality x>2x > 2 defines the region of the xyxy plane to the right of the line. Similarly x<2x < 2 defines the region to its left.

12345−2−1123xyx > 2
Figure 3.20. The region x>2x > 2, with its boundary line drawn broken.

Remark.

The region defined by x>2x > 2 does not include the line x=2x = 2. When a region does not include the points on its boundary lines, these are drawn as broken lines; when it does include them, they are drawn as solid lines.

Now consider the function f(x)=(x−3)(x+1)f(x) = (x - 3)(x + 1), and the curve representing it drawn on the xyxy plane. If PP, QQ and RR are points on a line x=x1x = x_1, with QQ on the curve, the yy-coordinate of QQ is f(x1)f(x_1). The yy-coordinate of PP, above QQ, is greater than that of QQ, and the yy-coordinate of RR, below QQ, is less. This argument applies for all values of x1x_1. Therefore the inequality y>(x−3)(x+1)y > (x - 3)(x + 1) defines the set of points in the region above the curve, and the inequality y<(x−3)(x+1)y < (x - 3)(x + 1) defines the region below it, whereas only for points on the curve is y=(x−3)(x+1)y = (x - 3)(x + 1). This last is called the equation of the curve, and the curve is often referred to simply as the curve y=(x−3)(x+1)y = (x - 3)(x + 1).

−2−11234−4−2246xyPQR
Figure 3.21. The region y>(x−3)(x+1)y > (x - 3)(x + 1), with the points PP, QQ, RR on one vertical line.

An equation such as x2−7x+3=0x^2 - 7x + 3 = 0 contains only one variable, and its solution comprises a finite set of values of xx. An equation containing two variables, such as y=(x−3)(x+1)y = (x - 3)(x + 1), has as its solution an infinite set of ordered pairs (x,y)(x, y). If AA is the solution set of the equation y=f(x)y = f(x), the elements of AA are the coordinates (x,y)(x, y) of all points on the curve y=f(x)y = f(x), and conversely the coordinates of points not on the curve are not elements of AA. So for a point PP on the curve and a point QQ off it, P∈AP \in A but Q∉AQ \notin A.

In general, if f(x)f(x) is any function of xx, then in the xyxy plane

  1. y=f(x)y = f(x) defines the curve representing f(x)f(x), and is called the equation of that curve;
  2. y>f(x)y > f(x) and y<f(x)y < f(x) define the regions of the plane above and below that curve.

Example 3.20.

Determine whether the points (5,11)(5, 11) and (−2,−20)(-2, -20) are on the curve y=(x−4)(x+6)y = (x - 4)(x + 6).

Substituting 1111 for yy in the left-hand side of the equation gives 1111, and substituting 55 for xx in the right-hand side gives (5−4)(5+6)=11(5 - 4)(5 + 6) = 11. The two sides are equal, so (5,11)(5, 11) is a member of the solution set of y=(x−4)(x+6)y = (x - 4)(x + 6) and is on the curve.

Substituting −20-20 for yy in the left-hand side gives −20-20, and substituting −2-2 for xx in the right-hand side gives (−2−4)(−2+6)=−24(-2 - 4)(-2 + 6) = -24. So the left-hand side is greater than the right-hand side, and (−2,−20)(-2, -20) is not a point on the curve: it lies above it.

Example 3.21.

Draw a sketch to show the region of the xyxy plane defined by the inequalities 0⩽x⩽20 \leqslant x \leqslant 2, y⩾0y \geqslant 0, y⩽x2y \leqslant x^2.

The relationship 0⩽x⩽20 \leqslant x \leqslant 2 contains two inequalities, x⩾0x \geqslant 0 and x⩽2x \leqslant 2, which must be considered separately. Taking each inequality in turn, we shade out the region that it excludes.

  1. y⩾0y \geqslant 0 is the line y=0y = 0 and the region above the xx-axis, so we shade out the region below the xx-axis.
  2. y⩽x2y \leqslant x^2 is the curve y=x2y = x^2 and the region below it, so we shade out the region above the curve.
  3. x⩾0x \geqslant 0 is the yy-axis and the region to its right, so we shade out the region to the left of the yy-axis.
  4. x⩽2x \leqslant 2 is the line x=2x = 2 and the region to its left, so we shade out the region to the right of this line.

Combining these four, the unshaded region, including the boundary lines, is the set of points that satisfies all the given inequalities.

121234xyy = x²x = 2
Figure 3.22. The region 0⩽x⩽20 \leqslant x \leqslant 2, y⩾0y \geqslant 0, y⩽x2y \leqslant x^2. Here the wanted region is shaded, rather than the regions excluded.

The Equation of a Particular Curve

So far in our work on functions and graphs we have begun with a function and deduced from its properties the curve that represents it. The reverse process, in which we begin with a curve given geometrically and deduce its equation, is as follows.

Consider the circle whose centre is the point C(4,2)C(4, 2) and whose radius is 22. Any point PP on the circumference of this circle is such that PC=2PC = 2; any point QQ inside the circle satisfies CQ<2CQ < 2; and any point RR outside the circle satisfies CR>2CR > 2. The distance of any point (x,y)(x, y) from CC is (x−4)2+(y−2)2\sqrt{(x - 4)^2 + (y - 2)^2}, so the coordinates (x,y)(x, y) of PP must satisfy the equation

(x−4)2+(y−2)2=4.(x - 4)^2 + (y - 2)^2 = 4 .
2461234xyCPQR
Figure 3.23. The circle with centre C(4,2)C(4, 2) and radius 22, with PP on it, QQ inside and RR outside.

This equation defines the set of points on the circumference of the circle, and so is the equation of the circle. Similarly the coordinates of QQ satisfy the inequality (x−4)2+(y−2)2<4(x - 4)^2 + (y - 2)^2 < 4, so this inequality defines the region inside the circle, and the inequality (x−4)2+(y−2)2>4(x - 4)^2 + (y - 2)^2 > 4 defines the region outside it.

The Straight Line

Straight lines play an important part in any geometric analysis. A straight line may be defined in many ways, for example as

  1. the line which passes through the origin and has a gradient of 12\tfrac12, or
  2. the line which passes through the points (2,1)(2, 1) and (−4,−2)(-4, -2).

For the first, if P(x,y)P(x, y) is a point on the line other than the origin, then the gradient of OPOP is 12\tfrac12. The gradient of OPOP is y−0x−0=yx\dfrac{y - 0}{x - 0} = \dfrac{y}{x}, so the coordinates of PP satisfy

yx=12,or2y=x,\frac{y}{x} = \frac12, \qquad\text{or}\qquad 2y = x,

and 2y=x2y = x is the equation of the line.

Remark.

For any point Q(x,y)Q(x, y) above PP, y>12xy > \tfrac12 x, that is 2y>x2y > x. Therefore the inequality 2y>x2y > x defines the region above the line, and similarly 2y<x2y < x defines the region below it.

For the second, a point P(x,y)P(x, y) is on the line through A(2,1)A(2, 1) and B(−4,−2)B(-4, -2) exactly when the gradient of PAPA equals the gradient of ABAB. The gradient of PAPA is y−1x−2\dfrac{y - 1}{x - 2} and the gradient of ABAB is 1−(−2)2−(−4)=12\dfrac{1 - (-2)}{2 - (-4)} = \dfrac12, so the coordinates of PP satisfy

y−1x−2=12,or2y=x.\frac{y - 1}{x - 2} = \frac12, \qquad\text{or}\qquad 2y = x .

These apparently different definitions give the same line. It is conventional to use integers for coefficients whenever possible.

Consider the more general case of the line whose gradient is mm and which passes through the origin. For a point P(x,y)P(x, y) on this line other than the origin, the gradient of OPOP is mm, so the coordinates of PP satisfy yx=m\dfrac{y}{x} = m, or y=mxy = mx; the final equation also includes the origin.

Generalising even further to cover any straight line, consider the line whose gradient is mm and which cuts the yy-axis at a directed distance cc from the origin. The number cc is called the intercept on the yy-axis. With AA the point (0,c)(0, c), a point P(x,y)P(x, y) is on the line exactly when the gradient of APAP is mm, so the coordinates of PP satisfy

y−cx−0=m,ory=mx+c.\frac{y - c}{x - 0} = m, \qquad\text{or}\qquad y = mx + c .
xycP(x, y)gradient m
Figure 3.24. The line y=mx+cy = mx + c, with intercept cc on the yy-axis and gradient mm.

This is called the standard form of the equation of a straight line. It follows that

  1. an equation of the form y=mx+cy = mx + c represents a straight line with gradient mm and intercept cc on the yy-axis;
  2. any equation involving a linear relationship between xx and yy, that is ax+by+c=0ax + by + c = 0 where aa and bb are constants not both zero, is the equation of a straight line.

Example 3.22.

Write down the gradient of the line 3x−4y+2=03x - 4y + 2 = 0, and find the equation of the line through the origin which is perpendicular to the given line.

Writing 3x−4y+2=03x - 4y + 2 = 0 in standard form gives y=34x+12y = \tfrac34 x + \tfrac12, so the gradient of the given line is 34\tfrac34. So the gradient of the perpendicular line is −43-\tfrac43. The required line passes through the origin, that is it has zero intercept on the yy-axis, so its equation is

y=−43x,that is4x+3y=0.y = -\tfrac43 x, \qquad\text{that is}\qquad 4x + 3y = 0 .

Example 3.23.

Sketch the line x−2y+3=0x - 2y + 3 = 0.

This line can be located accurately in the xyxy plane once two points on the line are known. The intercepts on the axes can be found by inspection: x=0x = 0 gives y=32y = \tfrac32, and y=0y = 0 gives x=−3x = -3. So the line passes through (0,32)\left(0, \tfrac32\right) and (−3,0)(-3, 0).

The Line with Gradient mm Through the Point (x1,y1)(x_1, y_1)

If P(x,y)P(x, y) is any point on the line with gradient mm passing through A(x1,y1)A(x_1, y_1), then the gradient of APAP is mm. Therefore the coordinates of PP satisfy

y−y1x−x1=m,that isy−y1=m(x−x1).\frac{y - y_1}{x - x_1} = m, \qquad\text{that is}\qquad y - y_1 = m(x - x_1) .

Example 3.24.

Find the equation of the line with gradient −13-\tfrac13 passing through (2,−1)(2, -1).

Substituting −13-\tfrac13 for mm, 22 for x1x_1 and −1-1 for y1y_1 gives the equation of the line as

y−(−1)=−13(x−2),that isx+3y+1=0.y - (-1) = -\tfrac13(x - 2), \qquad\text{that is}\qquad x + 3y + 1 = 0 .

Alternatively, as any straight line has an equation y=mx+cy = mx + c, the equation of this line can be written y=−13x+cy = -\tfrac13 x + c. As the point (2,−1)(2, -1) lies on the line, its coordinates satisfy the equation, so −1=−23+c-1 = -\tfrac23 + c and c=−13c = -\tfrac13. Therefore the equation is y=−13x−13y = -\tfrac13 x - \tfrac13, or x+3y+1=0x + 3y + 1 = 0.

Remark.

The worked examples necessarily contain a lot of explanation, but this should not mislead the reader into thinking that solutions need be equally long. The temptation to overwork a problem should be avoided, particularly in coordinate geometry, where problems are basically simple. With a little practice, either of the methods above gives the equation of a line directly.

The Line Through (x1,y1)(x_1, y_1) and (x2,y2)(x_2, y_2)

When x1≠x2x_1\neq x_2, the gradient of the line through (x1,y1)(x_1, y_1) and (x2,y2)(x_2, y_2) is y2−y1x2−x1\dfrac{y_2 - y_1}{x_2 - x_1}, so by the previous section its equation is

y−y1=y2−y1x2−x1(x−x1).y - y_1 = \frac{y_2 - y_1}{x_2 - x_1}(x - x_1) .

For example, the line through (1,−2)(1, -2) and (3,5)(3, 5) has equation

y+2=5−(−2)3−1(x−1),that is7x−2y−11=0.y + 2 = \frac{5 - (-2)}{3 - 1}(x - 1), \qquad\text{that is}\qquad 7x - 2y - 11 = 0 .

Intersection

If two curves cut at a point AA, then AA is called a point of intersection of the curves. The coordinates of AA satisfy both equations, so AA can be found by solving the equations simultaneously.

Example 3.25.

Find the point of intersection of the lines y−3x+1=0y - 3x + 1 = 0 and y+x−2=0y + x - 2 = 0.

Subtracting the first equation from the second gives 4x−3=04x - 3 = 0, so x=34x = \tfrac34, and then y=2−x=54y = 2 - x = \tfrac54. Therefore (34,54)\left(\tfrac34, \tfrac54\right) is the point of intersection.

Example 3.26.

Find the points of intersection AA and BB of the circle x2+y2−3x+2=0x^2 + y^2 - 3x + 2 = 0 and the line y=x−1y = x - 1.

The coordinates of both AA and BB satisfy both equations. Solving them simultaneously by substituting x−1x - 1 for yy in the equation of the circle,

x2+(x−1)2−3x+2=0,that is2x2−5x+3=0,that is(2x−3)(x−1)=0,x^2 + (x - 1)^2 - 3x + 2 = 0, \qquad\text{that is}\qquad 2x^2 - 5x + 3 = 0, \qquad\text{that is}\qquad (2x - 3)(x - 1) = 0,

so x=32x = \tfrac32 or x=1x = 1. Substituting these in y=x−1y = x - 1 gives y=12y = \tfrac12 and y=0y = 0. Therefore AA and BB are the points (32,12)\left(\tfrac32, \tfrac12\right) and (1,0)(1, 0).

In general, the coordinates of the points of intersection of two curves y=f(x)y = f(x) and y=g(x)y = g(x) can be found from the simultaneous solution of the equations y=f(x)y = f(x) and y=g(x)y = g(x).

Example 3.27.

Find the equation of the line through (1,2)(1, 2) which is perpendicular to the line 3x−7y+2=03x - 7y + 2 = 0.

Writing 3x−7y+2=03x - 7y + 2 = 0 in standard form gives y=37x+27y = \tfrac37 x + \tfrac27, showing that the given line has a gradient of 37\tfrac37. So the required line has gradient −73-\tfrac73 and passes through (1,2)(1, 2). Using y−y1=m(x−x1)y - y_1 = m(x - x_1) gives its equation as

y−2=−73(x−1),that is7x+3y−13=0.y - 2 = -\tfrac73(x - 1), \qquad\text{that is}\qquad 7x + 3y - 13 = 0 .

Note that the line perpendicular to 3x−7y+2=03x - 7y + 2 = 0 has an equation 7x+3y−13=07x + 3y - 13 = 0: the coefficients of xx and yy have been transposed, and the sign between the xx and yy terms has changed. In fact, given the line ax+by+c=0ax + by + c = 0, any line perpendicular to it has an equation

bx−ay+k=0,bx - ay + k = 0,

and this property of perpendicular lines can be used to shorten the working of problems.

Example 3.28 (The circumcentre of a triangle).

AA, BB and CC are the points (0,4)(0, 4), (2,3)(2, 3) and (−2,−1)(-2, -1). Find the circumcentre of triangle ABCABC.

The circumcentre of a triangle is the point of intersection of the perpendicular bisectors of its sides.

ACAC has gradient 4−(−1)0−(−2)=52\dfrac{4 - (-1)}{0 - (-2)} = \dfrac52, and its midpoint is (−1,32)\left(-1, \tfrac32\right). Therefore the perpendicular bisector of ACAC has gradient −25-\tfrac25 and passes through (−1,32)\left(-1, \tfrac32\right), so its equation is

y−32=−25(x+1),that is4x+10y−11=0.(1)y - \tfrac32 = -\tfrac25(x + 1), \qquad\text{that is}\qquad 4x + 10y - 11 = 0 . \tag{1}

Similarly the gradient of ABAB is 3−42−0=−12\dfrac{3 - 4}{2 - 0} = -\dfrac12, and its midpoint is (1,72)\left(1, \tfrac72\right). Therefore the perpendicular bisector of ABAB has gradient 22 and passes through (1,72)\left(1, \tfrac72\right), so its equation is

y−72=2(x−1),that is4x−2y+3=0.(2)y - \tfrac72 = 2(x - 1), \qquad\text{that is}\qquad 4x - 2y + 3 = 0 . \tag{2}

Subtracting (2)(2) from (1)(1) gives 12y−14=012y - 14 = 0, so y=76y = \tfrac76, and then 4x=2y−3=−234x = 2y - 3 = -\tfrac23 gives x=−16x = -\tfrac16. Therefore the circumcentre of triangle ABCABC is the point (−16,76)\left(-\tfrac16, \tfrac76\right).

Problem 3.10.

  1. Find the equation of the line passing through (−1,3)(-1, 3) and (−4,−3)(-4, -3).
  2. Find the equation of the line through (1,−2)(1, -2) perpendicular to 2x−3y+6=02x - 3y + 6 = 0.
  3. Draw a sketch showing the region of the xyxy plane defined by y<2y < 2 and y>(x−2)(x+2)y > (x - 2)(x + 2).

Remark (Summary).

If AA and BB are the points (x1,y1)(x_1, y_1) and (x2,y2)(x_2, y_2), then

  1. the length of ABAB is (x2−x1)2+(y2−y1)2\sqrt{(x_2 - x_1)^2 + (y_2 - y_1)^2};
  2. the midpoint of ABAB is the point (12(x1+x2),12(y1+y2))\left(\tfrac12(x_1 + x_2), \tfrac12(y_1 + y_2)\right);
  3. when x1≠x2x_1\neq x_2, the gradient of ABAB is y2−y1x2−x1\dfrac{y_2 - y_1}{x_2 - x_1}. A vertical line has no finite gradient and has an equation of the form x=constantx=\text{constant}.

The equation y=mx+cy = mx + c defines the straight line with gradient mm and intercept cc on the yy-axis. The inequality y>mx+cy > mx + c defines the region of the xyxy plane above that line, and y<mx+cy < mx + c the region below it.

If lines l1l_1 and l2l_2 have equations y=m1x+c1y = m_1 x + c_1 and y=m2x+c2y = m_2 x + c_2, then l1l_1 and l2l_2 are parallel if m1=m2m_1 = m_2, and, when both gradients are finite and non-zero, perpendicular if m1m2=−1m_1 m_2 = -1. A horizontal line is perpendicular to a vertical line. The equation of any line perpendicular to ax+by+c=0ax + by + c = 0 is of the form bx−ay+k=0bx - ay + k = 0.

Exercises

Questions marked with an examining board are taken from past A-level papers: JMB is the Joint Matriculation Board, U of L the University of London, C Cambridge, and AEB the Associated Examining Board.

Exercise 3.1.

In Utopia, assume income is non-negative. Income tax on earnings is calculated as follows: the first £10,000 is tax free, the next £10,000 is taxed at 5%5\%, and the remaining income is taxed at 10%10\%.

  1. Taking income as input and tax payable as output, state whether these rules for calculating tax constitute a function. If they do, state the implied domain and range.
  2. If £II is income and £TT is the tax payable, express the mapping I↦TI \mapsto T as formulae of the form T=f(I)T = f(I), stating the values of II for which each is valid.

Exercise 3.2.

State the inverse function, with its domain, of each of the following functions, or explain why it has none.

  1. f:x↦12x−3f : x \mapsto \tfrac12 x - 3, x∈Rx \in \RR;
  2. f:x↦(x−1)(x−3)f : x \mapsto (x - 1)(x - 3), x∈Rx \in \RR, x⩾2x \geqslant 2;
  3. f:x↦(x−2)(x+2)f : x \mapsto (x - 2)(x + 2), x∈Rx \in \RR;
  4. f:x↦10xf : x \mapsto 10^x, x∈Rx \in \RR. Draw sketch graphs of ff and f−1f^{-1} on the same set of axes, and describe how one graph can be obtained from the other.

Exercise 3.3.

Let f(x)=11−xf(x) = \dfrac{1}{1 - x}.

  1. For what value of xx is f(x)f(x) undefined? Describe the behaviour of f(x)f(x) as xx approaches this value from above and from below.
  2. Write down lim⁡x→∞f(x)\displaystyle\lim_{x \to \infty} f(x) and lim⁡x→−∞f(x)\displaystyle\lim_{x \to -\infty} f(x).
  3. Use this information to sketch the graph of ff, marking the asymptotes clearly.

Exercise 3.4.

Find the ranges of values of kk for which the equation x2+(k−3)x+k=0x^2 + (k - 3)x + k = 0 has

  1. real distinct roots;
  2. roots of the same sign. (JMB)

Exercise 3.5.

If xx is real and x2+(2−k)x+1−2k=0x^2 + (2 - k)x + 1 - 2k = 0, show that kk cannot lie between certain limits, and find these limits. (JMB)

Exercise 3.6.

Show that, if x2>k(x+1)x^2 > k(x + 1) for all real xx, then −4<k<0-4 < k < 0. (C)

Exercise 3.7.

Find the condition that must be satisfied by kk in order that the expression 2x2+6x+1+k(x2+2)2x^2 + 6x + 1 + k(x^2 + 2) may be positive for all real values of xx. (JMB)

Exercise 3.8.

  1. If x=2x = 2 is a root of the equation ax2+2(2a−5)x+8=0ax^2 + 2(2a - 5)x + 8 = 0, find the possible value, or values, of aa and the corresponding value, or values, of the other root.
  2. Find the range, or ranges, of possible values of the real number aa if ax2+2(2a−5)x+8>0ax^2 + 2(2a - 5)x + 8 > 0 for all real values of xx. (C)

Exercise 3.9.

Determine, for each of the expressions f(x)=x2+4x−6f(x) = x^2 + 4x - 6 and g(x)=−x2−8x+2g(x) = -x^2 - 8x + 2, the range, or ranges, of values of xx for which it is positive. Give your answers correct to two places of decimals, and explain briefly the reasons for your answers. (C)

Exercise 3.10.

  1. State the range of values of xx for which 2x2+5x−122x^2 + 5x - 12 is negative.
  2. The value of the constant aa is such that the quadratic function f(x)=x2+4x+a+3f(x) = x^2 + 4x + a + 3 is never negative. Determine the nature of the roots of the equation af(x)=(x+2)(a−1)a f(x) = (x + 2)(a - 1), and deduce the value of aa for which this equation has equal roots. (AEB, 1973)

Exercise 3.11.

By eliminating xx and yy from the equations

1x+1y=1,x+y=a,yx=m,\frac{1}{x} + \frac{1}{y} = 1, \qquad x + y = a, \qquad \frac{y}{x} = m,

where a≠0a \neq 0, obtain a relation between mm and aa. Given that aa is real, determine the ranges of values of aa for which mm is real. (JMB)

Exercise 3.12.

If a>0a > 0, show that the quadratic expression ax2+bx+cax^2 + bx + c is positive for all real values of xx when b2<4acb^2 < 4ac. Hence find the range of values of pp for which the quadratic function

f(x)=4x2+4px−(3p2+4p−3)f(x) = 4x^2 + 4px - (3p^2 + 4p - 3)

is positive for all real values of xx. Illustrate your result by making sketch graphs of f(x)f(x) for each of the cases p=0p = 0 and p=1p = 1. (U of L)

Exercise 3.13.

  1. If aa is a positive constant, find the set of values of xx for which a(x2+2x−8)a(x^2 + 2x - 8) is negative. Find the value of aa if this function has a least value of −27-27.
  2. Find two quadratic functions of xx which are zero at x=1x = 1, which take the value 1010 when x=0x = 0, and which have a greatest value of 1818. Sketch the graphs of these two functions. (U of L)

Exercise 3.14.

Find the set of values of kk for which f(x)=3x2−5x+kf(x) = 3x^2 - 5x + k is greater than unity for all real values of xx. Show that, for all kk, the least value of f(x)f(x) occurs when x=56x = \tfrac56, and find kk if this least value is zero. (U of L)

Exercise 3.15.

The roots of the equation 9x2+6x+1=4kx9x^2 + 6x + 1 = 4kx, where kk is a real constant, are denoted by α\alpha and β\beta.

  1. Show that the equation whose roots are 1α\dfrac{1}{\alpha} and 1β\dfrac{1}{\beta} is x2+6x+9=4kxx^2 + 6x + 9 = 4kx.
  2. Find the set of values of kk for which α\alpha and β\beta are real.
  3. Find also the set of values of kk for which α\alpha and β\beta are real and positive. (U of L)

Exercise 3.16.

  1. The points A(1,3)A(1, 3), B(5,7)B(5, 7), C(4,8)C(4, 8) and D(a,b)D(a, b) form a rectangle ABCDABCD. Find aa and bb.
  2. The points A(1,5)A(1, 5), B(4,−1)B(4, -1) and C(−2,−4)C(-2, -4) form triangle ABCABC. Show that the triangle is right-angled, and find its area.

Exercise 3.17.

ABCDABCD is a quadrilateral, where AA, BB, CC and DD are the points (3,−1)(3, -1), (6,0)(6, 0), (7,3)(7, 3) and (4,2)(4, 2). Show that the diagonals bisect each other at right angles, and hence find the area of ABCDABCD.

Exercise 3.18.

A circle of radius two units, with its centre at the origin, cuts the xx-axis at AA and BB and cuts the positive yy-axis at CC. Show that ABAB subtends a right angle at CC. If D(a,b)D(a, b) is a point on the circumference of the circle, find a relationship between aa and bb.

Exercise 3.19.

A point P(a,b)P(a, b) is equidistant from the yy-axis and from the point (4,0)(4, 0). Find a relationship between aa and bb.

Exercise 3.20.

Find the equation of the perpendicular from the point A(5,3)A(5, 3) to the line 2x−y+4=02x - y + 4 = 0. Hence find the distance of AA from the line.

Exercise 3.21.

The equation of a circle is (x−1)2+(y−1)2=4(x - 1)^2 + (y - 1)^2 = 4.

  1. Find the coordinates of AA and BB, the points of intersection of the line x+y=2x + y = 2 and the circle.
  2. Show that the point C(1,3)C(1, 3) is on the circumference of the circle. Find the midpoint MM of ABAB, and show that MC=MA=MBMC = MA = MB.
  3. What can you deduce about the line ABAB?

Exercise 3.22.

A line is drawn through the point A(1,2)A(1, 2) to cut the line 2y=3x−52y = 3x - 5 at PP and the line x+y=12x + y = 12 at QQ, with PP between AA and QQ. If AQ=2APAQ = 2AP, find the coordinates of PP and QQ. (U of L)

Check Yourself

 

Fresh questions on the whole chapter — none of them is worked out above. Do each on paper first; the box only tells you whether you got there.

Answers are checked in your browser, as often as you like. Nothing is sent anywhere and nothing is kept but your own work. A formula may be written with the symbols themselves or with ~ & | -> <-> ^, and \and, \or, \to expand as you type.

Exercise 3.23.

Each mapping below has domain x∈Rx \in \RR. Which one is one-one?

answer one of these

Exercise 3.24.

What is the range of f(x)=(x−3)2+2f(x) = (x - 3)^2 + 2, x∈Rx \in \RR?

answer one of these

Exercise 3.25.

What is the axis of symmetry of the curve y=x2−6x+1y = x^2 - 6x + 1?

answer one of these

Exercise 3.26.

What is the greatest value of 5−(x+1)25 - (x + 1)^2?

answer one of these

Exercise 3.27.

Which values of xx satisfy −2x<6-2x < 6?

answer one of these

Exercise 3.28.

Which values of xx satisfy (x+1)(x−3)<0(x + 1)(x - 3) < 0?

answer one of these

Exercise 3.29.

Which line is an asymptote to the curve y=3xy = 3^x?

answer one of these

Exercise 3.30.

Where does the curve y=1x−2y = \dfrac{1}{x - 2} have a discontinuity?

answer one of these

Exercise 3.31.

How is the curve y=3−x2y = 3 - x^2 obtained from the curve y=x2y = x^2?

answer one of these

Exercise 3.32.

What is the inverse of f(x)=3x−2f(x) = 3x - 2, x∈Rx \in \RR?

answer one of these

Exercise 3.33.

What is the length of the line joining (3,−4)(3, -4) to (−7,2)(-7, 2)?

answer one of these

Exercise 3.34.

What is the midpoint of the line joining (−1,−3)(-1, -3) to (3,−5)(3, -5)?

answer one of these

Exercise 3.35.

What is the gradient of a line perpendicular to the line joining (−1,5)(-1, 5) and (2,−3)(2, -3)?

answer one of these

Exercise 3.36.

What is the equation of the line through the origin perpendicular to 3x−2y+4=03x - 2y + 4 = 0?

answer one of these

Exercise 3.37.

What is the equation of the line with gradient 11 passing through the point (h,k)(h, k)?

answer one of these

Exercise 3.38.

At which points do the curves y=x2y = x^2 and y=x(2−x)y = x(2 - x) intersect?

answer one of these

Lesson 4

Limits, and Differentiation

Taught

Limits

The expression

(x+1)(x−2)x−2\frac{(x + 1)(x - 2)}{x - 2}

is not defined when x=2x = 2, where it would read 00\tfrac00. For every other value of xx the factor x−2x - 2 cancels, and the expression is simply x+1x + 1. Its graph is the line y=x+1y = x + 1 with one point missing: a hole at x=2x = 2.

−2−11234−112345xy(2, 3)
Figure 4.1. The graph of y=(x+1)(x−2)x−2y = \dfrac{(x + 1)(x - 2)}{x - 2}, which is the line y=x+1y = x + 1 with a hole at (2,3)(2, 3).

The expression is well defined for all values of xx other than 22, so we can legitimately ask what happens to it as xx approaches 22. Since it equals x+1x + 1 at all those values, it approaches 33. We say that its limit as xx approaches 22 is 33, and write

lim⁡x→2(x+1)(x−2)x−2=3.\lim_{x \to 2} \frac{(x + 1)(x - 2)}{x - 2} = 3 .

The value at x=2x = 2 itself plays no part. The limit is about the values of the expression near 22, and here there is no value at 22 at all.

Definition 4.1 (Limit).

Let f(x)f(x) be defined for all values of xx near aa, except possibly at aa itself. We say that f(x)f(x) tends to the limit kk as xx tends to aa, and write

lim⁡x→af(x)=k,\lim_{x \to a} f(x) = k,

if f(x)f(x) gets and stays as close to kk as we want, provided xx is sufficiently close to aa.

It is not enough for the values to keep getting closer to kk. The numbers 2.1,2.01,2.001,…2.1, 2.01, 2.001, \ldots get closer to 22, but each is also closer to 11 than the one before, and we want to say that they approach 22 and not 11. The difference is that they get as close to 22 as we want, however small we set the allowed distance, whereas every one of them stays more than 11 away from 11.

One-Sided Limits

Define f(x)f(x) to be x−1x - 1 when xx is negative and x+1x + 1 when xx is positive. Its graph is two half-lines with a jump between them at x=0x = 0.

−2−112−3−2−1123xy
Figure 4.2. The graph of f(x)f(x), equal to x−1x - 1 for negative xx and to x+1x + 1 for positive xx.

The limit of f(x)f(x) as x→0x \to 0 does not exist. We can make xx as close to 00 as we want, but f(x)f(x) does not get as close as we want to any one value: it is near −1-1 when xx is negative and near 11 when xx is positive. We met the same difficulty with 1x\dfrac1x in the last lesson, where the behaviour depended on whether xx approached zero from above or from below; here both sides settle down, but to different values. We write

lim⁡x→0+f(x)=1andlim⁡x→0−f(x)=−1,\lim_{x \to 0^+} f(x) = 1 \qquad\text{and}\qquad \lim_{x \to 0^-} f(x) = -1,

the first being the limit as xx tends to 00 from above, or from the right, and the second the limit from below, or from the left. The limit lim⁡x→af(x)\displaystyle\lim_{x \to a} f(x) exists only when these two one-sided limits exist and are equal.

Limits at Infinity

We can also take limits as xx tends to infinity. If g(x)=1xg(x) = \dfrac1x, then g(x)g(x) approaches 00 as xx grows, and we write

lim⁡x→∞g(x)=0.\lim_{x \to \infty} g(x) = 0 .

As xx tends to minus infinity g(x)g(x) also approaches 00, so lim⁡x→−∞g(x)=0\displaystyle\lim_{x \to -\infty} g(x) = 0 as well. For a limit at infinity, ”xx is sufficiently close to aa” in the definition of a limit is replaced by ”xx is large enough”, and for minus infinity by ”xx is negative and large enough in size”.

Problem 4.1.

Find each limit, or show that it does not exist.

  1. lim⁡x→1x2+x−2x−1\displaystyle\lim_{x \to 1} \frac{x^2 + x - 2}{x - 1};
  2. lim⁡x→1f(x)\displaystyle\lim_{x \to 1} f(x), where f(x)=x2f(x) = x^2 for x<1x < 1 and f(x)=2xf(x) = 2x for x>1x > 1;
  3. lim⁡x→∞2x+1x\displaystyle\lim_{x \to \infty} \frac{2x + 1}{x}.

The Gradient of a Curve

Chords, Tangents and Normals

Definition 4.2 (Chord, Tangent and Normal).

Let AA and BB be two points on a curve.

  1. The line joining AA and BB is a chord of the curve.
  2. The line touching the curve at AA is the tangent to the curve at AA.
  3. The line through AA perpendicular to the tangent at AA is the normal to the curve at AA.
xyABtangentnormalchord
Figure 4.3. A chord ABAB, and the tangent and normal at AA.

A straight line has the same gradient all along it. The gradient of a curve, which measures its slope, changes continually as we move along it. Suppose that as we move along the curve towards AA, the gradient stops changing at AA and stays constant from then on. We would then be moving along a straight line, and that line is the tangent at AA. So the gradient of the curve at AA is the same as the gradient of the tangent at AA.

Definition 4.3 (Gradient of a Curve).

The gradient of a curve at a point is the gradient of the tangent to the curve at that point. It measures the rate of increase of yy with respect to xx at that point.

If BB is another point on the curve, not too far from AA, the gradient of the chord ABAB is an approximate value for the gradient of the tangent at AA, and the closer BB is to AA the better the approximation.

An approximate value can also be found by plotting the curve, drawing the tangent by eye and measuring its gradient. This method has to be used when we know the coordinates of a finite number of points on a curve but not its equation, as with the data from an experiment. When the equation of the curve is known we want an accurate method, so that we can take the analysis of curves and functions further.

Approaching the Tangent

Take the parabola y=x2y = x^2 and the point (1.5,2.25)(1.5, 2.25) on it. The tangent there, drawn on the left of Figure 4.4, seems to rise about 33 units for each unit it moves to the right, so the gradient looks to be about 33. Is it exactly 33, and how would we find out?

11.52123456xythe tangent at (1.5, 2.25)1.41.51.61.722.252.52.75(1.6, 2.56)(1.5, 2.25)close up: the chord to (1.6, 2.56)
Figure 4.4. The tangent to y=x2y = x^2 at (1.5,2.25)(1.5, 2.25), and a close-up in which the dashed chord to (1.6,2.56)(1.6, 2.56) runs almost along the tangent.

To approximate it, take a second point on the parabola close to the first, such as (1.6,2.56)(1.6, 2.56), where 2.56=1.622.56 = 1.6^2 because the point lies on y=x2y = x^2. The line through (1.5,2.25)(1.5, 2.25) and (1.6,2.56)(1.6, 2.56) is found as in the last lesson. Its gradient is

2.56−2.251.6−1.5=0.310.1=3.1,\frac{2.56 - 2.25}{1.6 - 1.5} = \frac{0.31}{0.1} = 3.1,

and its equation is y=3.1x−2.4y = 3.1x - 2.4. So our approximation to the gradient is 3.13.1, and the close-up in Figure 4.4 shows how near the chord runs to the tangent.

The approximation cannot be made perfect by taking both points to be (1.5,2.25)(1.5, 2.25), since infinitely many lines pass through a single point and that tells us nothing. It can be made better. The chord through (1.5,2.25)(1.5, 2.25) and (1.501,2.253001)(1.501, 2.253001) has gradient

2.253001−2.251.501−1.5=0.0030010.001=3.001,\frac{2.253001 - 2.25}{1.501 - 1.5} = \frac{0.003001}{0.001} = 3.001,

so the gradients really do seem to be approaching 33. As BB approaches AA, the gradient of the chord ABAB approaches the gradient of the tangent at AA, that is

lim⁡B→A(gradient of chord AB)=gradient of tangent at A.\lim_{B \to A} \bigl(\text{gradient of chord } AB\bigr) = \text{gradient of tangent at } A .

Example 4.4.

Find the gradient of the curve y=x(2x−1)y = x(2x - 1) at the point AA where x=1x = 1.

When x=1x = 1, y=1y = 1, so AA is the point (1,1)(1, 1). We calculate the gradient of a chord ABAB and observe what happens to it as BB approaches AA. Take a succession of points B1,B2,B3,…B_1, B_2, B_3, \ldots where x=1.5,1.25,1.125,…x = 1.5, 1.25, 1.125, \ldots, each time halving the remaining difference between the xx-coordinates of AA and BB.

BBB1B_1B2B_2B3B_3B4B_4B5B_5
xx1.51.51.251.251.1251.1251.06251.06251.031251.03125
gradient of ABAB443.53.53.253.253.1253.1253.06253.0625

As BB approaches AA the gradient of the chord ABAB approaches 33, and we deduce that the gradient of the curve at AA is 33.

0.511.251.5123xyB₁B₂B₃Agradient 3
Figure 4.5. The chords AB1AB_1, AB2AB_2 and AB3AB_3 on y=x(2x−1)y = x(2x - 1), turning towards the tangent at A(1,1)A(1, 1).

The Delta Prefix

This numerical method is unsatisfactory, not least because of the amount of calculation involved, and in the end it only suggests the value 33. So instead of placing BB at particular positions we introduce a variable quantity for the difference between the xx-coordinates of AA and BB.

Definition 4.5 (Delta Prefix).

A variable quantity prefixed by δ\delta means a small increase in that quantity. So δx\delta x is a small increase in xx, and δy\delta y is a small increase in yy.

Remark.

The letter δ\delta is only a prefix. It cannot be treated as a factor: δx\delta x is one quantity, not δ\delta multiplied by xx.

Return to y=x(2x−1)y = x(2x - 1) and the point A(1,1)A(1, 1). If δx\delta x is the increase in the xx-coordinate in moving from AA to BB, then the xx-coordinate of BB is 1+δx1 + \delta x. Every point on the curve satisfies y=x(2x−1)y = x(2x - 1), so the yy-coordinate of BB is

(1+δx)[2(1+δx)−1]=(1+δx)(2 δx+1).(1 + \delta x)\bigl[2(1 + \delta x) - 1\bigr] = (1 + \delta x)(2\,\delta x + 1) .

The gradient of ABAB is the increase in yy divided by the increase in xx:

(1+δx)(2 δx+1)−1δx=3 δx+2(δx)2δx=3+2 δx.\frac{(1 + \delta x)(2\,\delta x + 1) - 1}{\delta x} = \frac{3\,\delta x + 2(\delta x)^2}{\delta x} = 3 + 2\,\delta x .

As BB approaches AA, the difference between their xx-coordinates approaches zero, that is δx→0\delta x \to 0. Therefore

gradient at A=lim⁡δx→0(3+2 δx)=3.\text{gradient at } A = \lim_{\delta x \to 0} \bigl(3 + 2\,\delta x\bigr) = 3 .

Dividing by δx\delta x is allowed because a limit only involves values of δx\delta x near zero, never zero itself.

The same calculation settles the question for the parabola. Writing hh for the small increase in xx from 1.51.5, the gradient of y=x2y = x^2 at (1.5,2.25)(1.5, 2.25) is

lim⁡h→0(1.5+h)2−1.52h=lim⁡h→01.52+3h+h2−1.52h=lim⁡h→03h+h2h=lim⁡h→0(3+h)=3,\lim_{h \to 0} \frac{(1.5 + h)^2 - 1.5^2}{h} = \lim_{h \to 0} \frac{1.5^2 + 3h + h^2 - 1.5^2}{h} = \lim_{h \to 0} \frac{3h + h^2}{h} = \lim_{h \to 0} \bigl(3 + h\bigr) = 3,

so it is exactly 33, as we guessed.

Example 4.6.

Find the gradient of the curve y=1xy = \dfrac1x at the point where x=2x = 2.

Let δx\delta x be the increase in xx in moving from A(2,12)A\left(2, \tfrac12\right) to a nearby point BB on the curve, so that BB is (2+δx,12+δx)\left(2 + \delta x, \dfrac{1}{2 + \delta x}\right). The gradient of the chord ABAB is

1δx(12+δx−12)=2−(2+δx)2(2+δx) δx=−δx2(2+δx) δx=−12(2+δx).\frac{1}{\delta x}\left(\frac{1}{2 + \delta x} - \frac12\right) = \frac{2 - (2 + \delta x)}{2(2 + \delta x)\,\delta x} = \frac{-\delta x}{2(2 + \delta x)\,\delta x} = -\frac{1}{2(2 + \delta x)} .

As δx→0\delta x \to 0, 2(2+δx)→42(2 + \delta x) \to 4, so the gradient at AA is −14-\tfrac14.

Problem 4.2.

Use the method of the last example to find the gradient of each curve at the point indicated.

  1. y=(x+1)(x−1)y = (x + 1)(x - 1), where x=2x = 2;
  2. y=2x(x−4)y = 2x(x - 4), where x=0x = 0;
  3. y=(x+2)(x−1)y = (x + 2)(x - 1), where x=−3x = -3.

The Gradient Function

We found that the gradient of y=x(2x−1)y = x(2x - 1) is 33 at the point where x=1x = 1. Instead of a fixed point, take AA to be any point (x,y)(x, y) on the curve. Its yy-coordinate can be written x(2x−1)x(2x - 1), since both coordinates satisfy the equation of the curve. Let BB be another point on the curve such that the increase in the xx-coordinate in moving from AA to BB is δx\delta x. Then BB is the point [x+δx,(x+δx)(2x+2 δx−1)]\bigl[x + \delta x, (x + \delta x)(2x + 2\,\delta x - 1)\bigr], and

gradient of chord AB=(x+δx)(2x+2 δx−1)−x(2x−1)δx=2x2+4x δx+2(δx)2−x−δx−2x2+xδx=4x δx−δx+2(δx)2δx=4x−1+2 δx.\begin{aligned} \text{gradient of chord } AB &= \frac{(x + \delta x)(2x + 2\,\delta x - 1) - x(2x - 1)}{\delta x} \\ &= \frac{2x^2 + 4x\,\delta x + 2(\delta x)^2 - x - \delta x - 2x^2 + x}{\delta x} \\ &= \frac{4x\,\delta x - \delta x + 2(\delta x)^2}{\delta x} \\ &= 4x - 1 + 2\,\delta x . \end{aligned}

Then the gradient at any point AA on the curve is

lim⁡δx→0(4x−1+2 δx)=4x−1.\lim_{\delta x \to 0} \bigl(4x - 1 + 2\,\delta x\bigr) = 4x - 1 .

So the function 4x−14x - 1 gives the gradient at every point on the curve y=x(2x−1)y = x(2x - 1). The gradient at a particular point, that is the rate of increase of yy with respect to xx there, is found by substituting the xx-coordinate of the point into 4x−14x - 1. At x=1x = 1 this gives 33, as before.

The function 4x−14x - 1 is called the gradient function of y=x(2x−1)y = x(2x - 1), and the process of deriving it is called differentiation with respect to xx. Since 4x−14x - 1 was derived from x(2x−1)x(2x - 1), it is also called the derivative, or derived function, of x(2x−1)x(2x - 1). Using ddx\dfrac{d}{dx} as a symbol for “the derivative with respect to xx of”, we write

ddx[x(2x−1)]=4x−1,\frac{d}{dx}\bigl[x(2x - 1)\bigr] = 4x - 1,

or, when y=x(2x−1)y = x(2x - 1),

dydx=4x−1.\frac{dy}{dx} = 4x - 1 .

Here dydx\dfrac{dy}{dx} means the derivative of yy with respect to xx, and is sometimes called the differential coefficient of yy. As an alternative notation we can use the symbol DD for “the derivative of”, so that D[x(2x−1)]=4x−1D\bigl[x(2x - 1)\bigr] = 4x - 1, and DD is referred to as the differential operator.

Since 4x−14x - 1 represents the rate of increase of yy with respect to xx, dydx\dfrac{dy}{dx} also represents the rate of increase of yy with respect to xx. Similarly dvdt\dfrac{dv}{dt} means the derivative of vv with respect to tt, or the rate of increase of vv with respect to tt.

Example 4.7.

Differentiate y=x3+3y = x^3 + 3 with respect to xx.

Let AA be the point (x,x3+3)(x, x^3 + 3) and BB the point [x+δx,(x+δx)3+3]\bigl[x + \delta x, (x + \delta x)^3 + 3\bigr] on y=x3+3y = x^3 + 3. Then

gradient of AB=(x+δx)3+3−(x3+3)δx=3x2 δx+3x(δx)2+(δx)3δx,\text{gradient of } AB = \frac{(x + \delta x)^3 + 3 - (x^3 + 3)}{\delta x} = \frac{3x^2\,\delta x + 3x(\delta x)^2 + (\delta x)^3}{\delta x},

which simplifies to 3x2+3x δx+(δx)23x^2 + 3x\,\delta x + (\delta x)^2. So the gradient at AA is lim⁡δx→0[3x2+3x δx+(δx)2]=3x2\displaystyle\lim_{\delta x \to 0} \bigl[3x^2 + 3x\,\delta x + (\delta x)^2\bigr] = 3x^2, that is

dydx=3x2.\frac{dy}{dx} = 3x^2 .

Problem 4.3.

Use the method of the last example to differentiate each of the following with respect to xx.

  1. y=x2y = x^2;
  2. y=x4y = x^4;
  3. y=5x2y = 5x^2;
  4. y=1xy = \dfrac1x.

Differentiating from First Principles

Consider any curve y=f(x)y = f(x). Let A[x,f(x)]A[x, f(x)] be any point on it, and let δx\delta x be the increase in the xx-coordinate in moving from AA to another point BB on the same curve, so that BB is the point [x+δx,f(x+δx)][x + \delta x, f(x + \delta x)]. Let δy\delta y be the corresponding increase in the yy-coordinate, that is

δy=f(x+δx)−f(x).\delta y = f(x + \delta x) - f(x) .
xyA[x, f(x)]B[x + δx, f(x + δx)]δxδy
Figure 4.6. The increases δx\delta x and δy\delta y in moving from AA to BB along y=f(x)y = f(x).

The gradient of the chord ABAB is

δyδx=f(x+δx)−f(x)δx,\frac{\delta y}{\delta x} = \frac{f(x + \delta x) - f(x)}{\delta x},

so the gradient at AA is

dydx=lim⁡δx→0δyδx=lim⁡δx→0f(x+δx)−f(x)δx.\frac{dy}{dx} = \lim_{\delta x \to 0} \frac{\delta y}{\delta x} = \lim_{\delta x \to 0} \frac{f(x + \delta x) - f(x)}{\delta x} .

Definition 4.8 (Derivative).

Let ff be a function and cc a number. The derivative of ff at cc is the gradient of the graph of ff at the point where x=cx = c, defined by

lim⁡h→0f(c+h)−f(c)h.\lim_{h \to 0} \frac{f(c + h) - f(c)}{h} .

A function for which this limit exists is differentiable at cc. The derivative of f(x)f(x) is written f′(x)f'(x) or ddxf(x)\dfrac{d}{dx} f(x), and when yy is a function of xx that we would graph, it is written dydx\dfrac{dy}{dx}.

The limit exists when the graph looks as though it has a slope at the point. If the graph had a spike there, it would not have a well-defined slope and the limit would not exist.

Using this definition to differentiate a function is called differentiating from first principles, and it can be used for any function, including functions we have not yet met. Fortunately it is not always necessary to go back to first principles, because whole families of functions can be differentiated by rules.

Remark.

We cannot tell how fast something is going from a single photograph of it. But we can photograph it at two nearby moments and divide the distance it moved by the time between the photographs, which gives a good approximation to its speed, and the closer the two moments, the better the approximation. This is how instruments that measure speed work, and it is the gradient of a chord standing in for the gradient of a tangent.

Rules of Differentiation

Differentiating a Constant

The equation y=cy = c, where cc is a constant, represents a straight line parallel to the xx-axis, so it has zero gradient:

ddx(c)=0.\frac{d}{dx}(c) = 0 .

Differentiating axax

The equation y=axy = ax, where aa is a constant, represents a straight line with gradient aa, so

ddx(ax)=a.\frac{d}{dx}(ax) = a .

Differentiating xnx^n

The table collects results from the examples and problems above.

f(x)f(x)xxx2x^2x3x^3x4x^4x−1x^{-1}
ddxf(x)\dfrac{d}{dx} f(x)112x2x3x23x^24x34x^3−x−2-x^{-2}

From this table it appears that to differentiate a power of xx we multiply by that power and then subtract one from the power, that is

ddx(xn)=nxn−1.(1)\frac{d}{dx}\left(x^n\right) = n x^{n-1} . \tag{1}

This is the power rule.

The Power Rule for Whole Numbers

If we replace 1.51.5 by xx and do the same calculation as for the parabola, the algebra gives

(x+h)2−x2=x2+2xh+h2−x2=2xh+h2,(x + h)^2 - x^2 = x^2 + 2xh + h^2 - x^2 = 2xh + h^2,

and after dividing by hh the remaining hh vanishes in the limit, so the parabola y=x2y = x^2 has gradient 2x2x at each value of xx. Expanding further powers shows a pattern:

(x+h)2−x2=2x h+h2,(x+h)3−x3=3x2h+3xh2+h3,(x+h)4−x4=4x3h+6x2h2+4xh3+h4.\begin{aligned} (x + h)^2 - x^2 &= 2x\,h + h^2, \\ (x + h)^3 - x^3 &= 3x^2 h + 3x h^2 + h^3, \\ (x + h)^4 - x^4 &= 4x^3 h + 6x^2 h^2 + 4x h^3 + h^4 . \end{aligned}

In each line the term with hh to the first power is nxn−1hn x^{n-1} h, and every other term contains h2h^2 or a higher power of hh. When nn is a positive whole number we show that this always happens by using the binomial expansion. The coefficients of (x+h)n(x + h)^n are those of (1+x)n(1 + x)^n, the numbers in row nn of Pascal’s triangle, so

(x+h)n=xn+(second number in row n) xn−1h+(terms containing h2,h3,…).(x + h)^n = x^n + (\text{second number in row } n)\, x^{n-1} h + (\text{terms containing } h^2, h^3, \ldots) .

The second number in a row is the sum of the two numbers above it, which are 11 and the second number of the row before. So it goes up by 11 from each row to the next, and as it is 11 in row one, it is nn in row nn. Therefore

(x+h)n−xnh=nxn−1+(terms containing h,h2,…),\frac{(x + h)^n - x^n}{h} = n x^{n-1} + (\text{terms containing } h, h^2, \ldots),

and as h→0h \to 0 this tends to nxn−1n x^{n-1}.

The power rule, deduced here for whole numbers, is valid for all powers of xx, including fractional and negative powers, although this cannot be shown at this stage, and for the time being we take it on trust. Under the power conventions in Lesson 1, if nn is not a whole number then we use xnx^n only where it is defined, and for the fractional powers considered here this means x>0x>0. The rule needs the power to be a fixed number: it does not find the gradient of xxx^x, where the power changes with xx.

Example 4.9.

Accepting that the power rule can be used for any power of xx,

ddx(x9)=9x8,ddx(x)=ddx(x1/2)=12x−1/2=12x(x>0),ddx(1x3)=ddx(x−3)=−3x−4=−3x4.\begin{aligned} \frac{d}{dx}\left(x^9\right) &= 9x^8, \\ \frac{d}{dx}\left(\sqrt{x}\right) &= \frac{d}{dx}\left(x^{1/2}\right) = \tfrac12 x^{-1/2} = \frac{1}{2\sqrt{x}} \quad (x>0), \\ \frac{d}{dx}\left(\frac{1}{x^3}\right) &= \frac{d}{dx}\left(x^{-3}\right) = -3x^{-4} = -\frac{3}{x^4} . \end{aligned}

Problem 4.4.

Differentiate each of the following with respect to xx by rule.

  1. x7x^7;
  2. 1x2\dfrac{1}{x^2};
  3. x3\sqrt{x^3};
  4. 1x3\dfrac{1}{\sqrt[3]{x}}.

Differentiating axnax^n

From the problems, ddx(5x2)=10x\dfrac{d}{dx}\left(5x^2\right) = 10x, which is 55 times the derivative of x2x^2. A constant factor always comes straight through, since

a(x+δx)n−axnδx=a×(x+δx)n−xnδx,\frac{a(x + \delta x)^n - a x^n}{\delta x} = a \times \frac{(x + \delta x)^n - x^n}{\delta x},

and so

ddx(axn)=anxn−1,(2)\frac{d}{dx}\left(a x^n\right) = a n x^{n-1}, \tag{2}

where aa is a constant.

Sums and Differences

From first principles,

(x+δx)2+3(x+δx)−(x2+3x)δx=2x+3+δx,\frac{(x + \delta x)^2 + 3(x + \delta x) - \left(x^2 + 3x\right)}{\delta x} = 2x + 3 + \delta x,

so ddx(x2+3x)=2x+3\dfrac{d}{dx}\left(x^2 + 3x\right) = 2x + 3, and in the same way ddx(x2−2x+1)=2x−2\dfrac{d}{dx}\left(x^2 - 2x + 1\right) = 2x - 2. Comparing these with the separate derivatives,

ddx(x2+3x)=2x+3=ddx(x2)+ddx(3x),ddx(x2−2x+1)=2x−2=ddx(x2)−ddx(2x)+ddx(1).\begin{aligned} \frac{d}{dx}\left(x^2 + 3x\right) &= 2x + 3 = \frac{d}{dx}\left(x^2\right) + \frac{d}{dx}(3x), \\ \frac{d}{dx}\left(x^2 - 2x + 1\right) &= 2x - 2 = \frac{d}{dx}\left(x^2\right) - \frac{d}{dx}(2x) + \frac{d}{dx}(1) . \end{aligned}

So the operation “differentiate” is distributive across the addition and subtraction of functions, and gradients add:

ddx[f(x)±g(x)]=ddxf(x)±ddxg(x).(3)\frac{d}{dx}\bigl[f(x) \pm g(x)\bigr] = \frac{d}{dx} f(x) \pm \frac{d}{dx} g(x) . \tag{3}

This holds in general, because the gradient of a chord of y=f(x)+g(x)y = f(x) + g(x) is the gradient of the corresponding chord of y=f(x)y = f(x) plus that of y=g(x)y = g(x).

Example 4.10.

Using rules (2)(2) and (3)(3),

ddx(3x2−7x)=3(2x)−7=6x−7,ddx(x+1x)=ddx(x+x−1)=1−x−2=1−1x2,ddx(x3−2x2)=3x2−4x.\begin{aligned} \frac{d}{dx}\left(3x^2 - 7x\right) &= 3(2x) - 7 = 6x - 7, \\ \frac{d}{dx}\left(x + \frac1x\right) &= \frac{d}{dx}\left(x + x^{-1}\right) = 1 - x^{-2} = 1 - \frac{1}{x^2}, \\ \frac{d}{dx}\left(x^3 - 2x^2\right) &= 3x^2 - 4x . \end{aligned}

Differentiation is not distributive across multiplication and division. For example

ddx[x(x−3)]=ddx(x2−3x)=2x−3,butddx(x)×ddx(x−3)=1×1=1.\frac{d}{dx}\bigl[x(x - 3)\bigr] = \frac{d}{dx}\left(x^2 - 3x\right) = 2x - 3, \qquad\text{but}\qquad \frac{d}{dx}(x) \times \frac{d}{dx}(x - 3) = 1 \times 1 = 1 .

Gradients do not multiply in the way they add. So in order to differentiate at this stage, any product must be expanded and any quotient must be divided out, to give terms which are added or subtracted, and no rule should be assumed that has not been established.

Example 4.11.

Differentiate f(x)=4x2+x−12xf(x) = \dfrac{4x^2 + x - 1}{2x}.

f(x)=4x22x+x2x−12x=2x+12−12x−1,f(x) = \frac{4x^2}{2x} + \frac{x}{2x} - \frac{1}{2x} = 2x + \frac12 - \frac12 x^{-1},

therefore

f′(x)=2+0−12(−x−2)=2+12x2.f'(x) = 2 + 0 - \tfrac12\left(-x^{-2}\right) = 2 + \frac{1}{2x^2} .

Example 4.12.

Find the gradient of the curve y=(x−3)(x2+2)y = (x - 3)\left(x^2 + 2\right) at the point on the curve where x=1x = 1.

Expanding, y=x3−3x2+2x−6y = x^3 - 3x^2 + 2x - 6, therefore

dydx=3x2−6x+2.\frac{dy}{dx} = 3x^2 - 6x + 2 .

When x=1x = 1, dydx=3−6+2=−1\dfrac{dy}{dx} = 3 - 6 + 2 = -1, so the gradient of y=(x−3)(x2+2)y = (x - 3)\left(x^2 + 2\right) is −1-1 at the point where x=1x = 1.

Example 4.13.

Find the coordinates of the point on the curve y=2x2y = \dfrac{2}{x^2} at which its gradient is 12\tfrac12.

Since y=2x−2y = 2x^{-2}, dydx=−4x−3\dfrac{dy}{dx} = -4x^{-3}. The gradient is 12\tfrac12 when

−4x3=12,that isx3=−8,x=−2.-\frac{4}{x^3} = \frac12, \qquad\text{that is}\qquad x^3 = -8, \qquad x = -2 .

When x=−2x = -2, y=24=12y = \dfrac{2}{4} = \tfrac12. Therefore the gradient of y=2x2y = \dfrac{2}{x^2} is 12\tfrac12 at the point (−2,12)\left(-2, \tfrac12\right).

Problem 4.5.

  1. Differentiate (x−3)(2x+5)(x - 3)(2x + 5) and x (x−1)\sqrt{x}\,(x - 1) with respect to xx.
  2. Find the gradient of the curve y=(2x−3)(x+1)y = (2x - 3)(x + 1) at the point where x=0x = 0.
  3. Find the coordinates of the point on the curve y=xy = \sqrt{x} at which the gradient is 22.

Tangents and Normals

The tangent to a curve at a point passes through that point and has the gradient of the curve there, so its equation comes from the line with a given gradient through a given point. When the tangent has non-zero gradient mm, the normal passes through the same point with the perpendicular gradient −1m-\dfrac1m. A horizontal tangent has a vertical normal.

Take y=x2y = x^2 at (1.5,2.25)(1.5, 2.25) once more. The tangent is the line of gradient 33 through (1.5,2.25)(1.5, 2.25), which is

y−2.25=3(x−1.5).y - 2.25 = 3(x - 1.5) .

Rearranging, 4y−12x+9=04y - 12x + 9 = 0 is the equation of the tangent.

The normal to y=x2y = x^2 at (1.5,2.25)(1.5, 2.25) is the line through that point perpendicular to the tangent, so its gradient is −13-\tfrac13 and its equation is

y−2.25=−13(x−1.5),that isy+13x−114=0.y - 2.25 = -\tfrac13(x - 1.5), \qquad\text{that is}\qquad y + \tfrac13 x - \tfrac{11}{4} = 0 .

Multiplying through by 1212 to clear the fractions gives 12y+4x−33=012y + 4x - 33 = 0.

−112312345xytangentnormal
Figure 4.7. The parabola y=x2y = x^2 with its tangent 4y−12x+9=04y - 12x + 9 = 0 and its normal 12y+4x−33=012y + 4x - 33 = 0 at (1.5,2.25)(1.5, 2.25).

Example 4.14.

Find the equation of the tangent to the curve y=x2−3x+2y = x^2 - 3x + 2 at the point where it cuts the yy-axis.

The curve cuts the yy-axis where x=0x = 0, at (0,2)(0, 2). The gradient of the tangent there is the value of dydx\dfrac{dy}{dx} when x=0x = 0. As y=x2−3x+2y = x^2 - 3x + 2, dydx=2x−3\dfrac{dy}{dx} = 2x - 3, so when x=0x = 0 the gradient of the curve is −3-3.

Therefore the tangent has gradient −3-3 and passes through (0,2)(0, 2). Its equation is y=−3x+2y = -3x + 2, that is

3x+y−2=0.3x + y - 2 = 0 .

Example 4.15.

Find the equation of the normal to the curve y=1xy = \dfrac1x at the point where x=2x = 2, and the coordinates of the point where this normal cuts the curve again.

As y=x−1y = x^{-1}, dydx=−1x2\dfrac{dy}{dx} = -\dfrac{1}{x^2}. So when x=2x = 2, y=12y = \tfrac12 and dydx=−14\dfrac{dy}{dx} = -\tfrac14. Therefore the tangent at (2,12)\left(2, \tfrac12\right) has gradient −14-\tfrac14, and the normal, the line perpendicular to the tangent, has gradient 44. The equation of the normal is

y−12=4(x−2),that is8x−2y−15=0.y - \tfrac12 = 4(x - 2), \qquad\text{that is}\qquad 8x - 2y - 15 = 0 .

The points of intersection of the normal and the curve are given by solving 8x−2y−15=08x - 2y - 15 = 0 and y=1xy = \dfrac1x simultaneously. Substituting for yy,

8x−2x−15=0,that is8x2−15x−2=0.8x - \frac2x - 15 = 0, \qquad\text{that is}\qquad 8x^2 - 15x - 2 = 0 .

We already know that the normal and the curve meet where x=2x = 2, so x−2x - 2 is a factor, and

(x−2)(8x+1)=0.(x - 2)(8x + 1) = 0 .

Therefore they meet again where x=−18x = -\tfrac18 and y=−8y = -8, at the point (−18,−8)\left(-\tfrac18, -8\right).

−1123−6−4−22xy(2, ½)(−⅛, −8)
Figure 4.8. The normal to y=1xy = \dfrac1x at (2,12)\left(2, \tfrac12\right) meets the other branch of the curve at (−18,−8)\left(-\tfrac18, -8\right).

Problem 4.6.

  1. Find the equations of the tangent and the normal to the curve y=x2−5x+2y = x^2 - 5x + 2 at the point where x=3x = 3.
  2. Find the equation of the tangent to y=2x2−3xy = 2x^2 - 3x which has a gradient of 11.
  3. Find the equation of the tangent to y=(x−5)(2x+1)y = (x - 5)(2x + 1) which is parallel to the xx-axis.

Stationary Values and Turning Points

Stationary Values

Definition 4.16 (Stationary Value).

A stationary value of a function f(x)f(x) is a value of f(x)f(x) at which its rate of change with respect to xx is zero, that is, where

ddxf(x)=0.\frac{d}{dx} f(x) = 0 .

A point on the curve y=f(x)y = f(x) at which dydx=0\dfrac{dy}{dx} = 0 is a stationary point.

At a stationary value of f(x)f(x) the gradient of the curve y=f(x)y = f(x) is zero, so the tangent to the curve is parallel to the xx-axis. The stationary values of f(x)f(x) are the yy-coordinates of the points on y=f(x)y = f(x) at which the tangent is parallel to the xx-axis.

The curve y=x3−2x2y = x^3 - 2x^2 in Figure 4.9 goes flat and turns round in two places, and the derivative tells us exactly where. We found that ddx(x3−2x2)=3x2−4x\dfrac{d}{dx}\left(x^3 - 2x^2\right) = 3x^2 - 4x, and the graph is flat where

3x2−4x=0,that isx(3x−4)=0.3x^2 - 4x = 0, \qquad\text{that is}\qquad x(3x - 4) = 0 .

The solutions are x=0x = 0 and x=43x = \tfrac43, so it is exactly at these two values of xx that the graph is flat and turns round, at (0,0)(0, 0) and (43,−3227)\left(\tfrac43, -\tfrac{32}{27}\right). A zero derivative does not always mean that the graph turns round, as we shall see with y=x3y = x^3.

−112−2−112xy(4/3, −32/27)
Figure 4.9. The curve y=x3−2x2y = x^3 - 2x^2, with horizontal tangents at (0,0)(0, 0) and (43,−3227)\left(\tfrac43, -\tfrac{32}{27}\right).

Example 4.17.

Find the stationary values of x3−3x2+2x^3 - 3x^2 + 2.

If f(x)=x3−3x2+2f(x) = x^3 - 3x^2 + 2, then ddxf(x)=3x2−6x\dfrac{d}{dx} f(x) = 3x^2 - 6x. At stationary values of f(x)f(x) this is zero, so

3x2−6x=0,3x(x−2)=0.3x^2 - 6x = 0, \qquad 3x(x - 2) = 0 .

So the stationary values of x3−3x2+2x^3 - 3x^2 + 2 occur when x=0x = 0 and when x=2x = 2. They are

f(0)=2andf(2)=23−3(22)+2=−2.f(0) = 2 \qquad\text{and}\qquad f(2) = 2^3 - 3\left(2^2\right) + 2 = -2 .

For the curve y=x3−3x2+2y = x^3 - 3x^2 + 2, the gradient is zero at the points (0,2)(0, 2) and (2,−2)(2, -2).

The derivative also tells us which way a function is going. Where dydx\dfrac{dy}{dx} is positive the curve is rising, and we say that yy is increasing; where dydx\dfrac{dy}{dx} is negative the curve is falling, and yy is decreasing. For y=x3−2x2y = x^3 - 2x^2 the derivative x(3x−4)x(3x - 4) is positive for x<0x < 0, negative for 0<x<430 < x < \tfrac43 and positive again for x>43x > \tfrac43, which is what Figure 4.9 shows.

Problem 4.7.

  1. Find the values of xx at which x3−12x+1x^3 - 12x + 1 has stationary values.
  2. Find the stationary value of 3x2−4x+23x^2 - 4x + 2.
  3. Find the coordinates of the point on y=(x−2)(x+3)y = (x - 2)(x + 3) at which the gradient is zero.

Turning Points

The gradient of a curve can be zero at several points. Close to one of these points, the shape of the curve belongs to one of the three kinds shown at AA, BB and CC in Figure 4.10.

xyABC
Figure 4.10. A maximum turning point AA, a minimum turning point BB, and a point CC where the gradient is zero but the curve does not turn.

Moving along the curve in the positive direction of the xx-axis:

  1. near AA the gradient changes from positive, through zero at AA, to negative;
  2. near BB the gradient changes from negative, through zero at BB, to positive;
  3. at CC the gradient is zero, but it does not change sign as we move through CC, so the curve does not turn at CC. What does change at CC is the sense in which the curve is turning, from clockwise to anticlockwise.

Definition 4.18 (Turning Points and Points of Inflexion).

  1. A point on a curve at which the gradient changes from positive, through zero, to negative is a maximum turning point. Its yy-coordinate is a maximum value of yy, or of f(x)f(x) where y=f(x)y = f(x).
  2. A point at which the gradient changes from negative, through zero, to positive is a minimum turning point. Its yy-coordinate is a minimum value of yy, or of f(x)f(x).
  3. A point on a curve at which the sense of turning changes is a point of inflexion.

The gradient at a maximum or minimum turning point must be zero. The terms maximum value and minimum value do not mean the same as greatest value and least value: maxima and minima describe the behaviour of a function only in the immediate neighbourhood of its stationary values. In Figure 4.10, to the left of AA, the curve falls below the minimum value at BB.

Apart from CC there are two other points of inflexion in Figure 4.10, one between AA and BB and another between BB and CC, so the gradient at a point of inflexion is not necessarily zero. A stationary point can therefore be a maximum, a minimum, or neither. The simplest example of the last kind is y=x3y = x^3 at the origin: its gradient 3x23x^2 is zero there, but the curve rises on both sides.

−11−2−112xy
Figure 4.11. The curve y=x3y = x^3, stationary at the origin without turning.

The Nature of a Stationary Value

We know how to find the points on y=f(x)y = f(x) at which f(x)f(x) has stationary values, but to tell them apart we need to look further, and there are several ways of doing this. In Figure 4.10, let A1A_1 and A2A_2 be points on the curve close to AA and to its left and right respectively, and let B1B_1, B2B_2 and C1C_1, C2C_2 be placed in the same way about BB and CC.

Values of yy

At AA, a maximum, yy at A1A_1 and yy at A2A_2 are both less than yy at AA. At BB, a minimum, yy at B1B_1 and yy at B2B_2 are both greater than yy at BB. At CC, a point of inflexion, yy at C1C_1 is less than yy at CC and yy at C2C_2 is greater.

MaximumMinimumInflexion
Values of yy either side of the stationary valueboth smallerboth largerone smaller and one larger

The Sign of the Gradient

At A1A_1 the gradient dydx\dfrac{dy}{dx} is positive, at AA it is zero and at A2A_2 it is negative. At B1B_1 it is negative, at BB zero and at B2B_2 positive. At C1C_1 it is positive, at CC zero and at C2C_2 positive again.

MaximumMinimumInflexion
Sign of dydx\dfrac{dy}{dx} moving through the stationary value+ 0 −+ \ 0 \ -− 0 +- \ 0 \ ++ 0 ++ \ 0 \ + or − 0 −- \ 0 \ -

The Second Derivative

When passing through AA, dydx\dfrac{dy}{dx} changes from positive to negative, so dydx\dfrac{dy}{dx} decreases as xx increases, that is, the rate of increase of dydx\dfrac{dy}{dx} with respect to xx is negative. The rate of increase of dydx\dfrac{dy}{dx} with respect to xx would be written ddx(dydx)\dfrac{d}{dx}\left(\dfrac{dy}{dx}\right), and this clumsy notation is condensed to

d2ydx2,\frac{d^2 y}{dx^2},

called the second derivative of yy with respect to xx. Similarly, when passing through BB, dydx\dfrac{dy}{dx} changes from negative to positive, so dydx\dfrac{dy}{dx} increases as xx increases and d2ydx2\dfrac{d^2 y}{dx^2} is positive.

MaximumMinimum
Sign of d2ydx2\dfrac{d^2 y}{dx^2}negative (or zero)positive (or zero)

Points of inflexion are not so easily dealt with by this method. In the smooth examples considered here, d2ydx2=0\dfrac{d^2 y}{dx^2} = 0 at such points, but d2ydx2\dfrac{d^2 y}{dx^2} can also be zero at maxima and minima.

The three tables summarise three alternative methods for determining the nature of a stationary value. The third method fails if d2ydx2\dfrac{d^2 y}{dx^2} is found to be zero at the stationary value, and in that case one of the first two has to be used.

Example 4.19.

Find the points on y=x4+4x3−6y = x^4 + 4x^3 - 6 at which the gradient is zero, and determine the nature of these points.

y=x4+4x3−6,dydx=4x3+12x2,d2ydx2=12x2+24x.y = x^4 + 4x^3 - 6, \qquad \frac{dy}{dx} = 4x^3 + 12x^2, \qquad \frac{d^2 y}{dx^2} = 12x^2 + 24x .

The gradient is zero when 4x3+12x2=04x^3 + 12x^2 = 0, that is 4x2(x+3)=04x^2(x + 3) = 0, so when x=−3x = -3 and when x=0x = 0.

When x=−3x = -3, d2ydx2=12(9)+24(−3)=36\dfrac{d^2 y}{dx^2} = 12(9) + 24(-3) = 36, which is positive, so yy has a minimum value here. It is y=(−3)4+4(−3)3−6=−33y = (-3)^4 + 4(-3)^3 - 6 = -33, so (−3,−33)(-3, -33) is a minimum turning point.

When x=0x = 0, y=−6y = -6 and d2ydx2=0\dfrac{d^2 y}{dx^2} = 0, which is inconclusive. So we look at the sign of dydx=4x2(x+3)\dfrac{dy}{dx} = 4x^2(x + 3) on either side of the point where x=0x = 0.

xx−12-\tfrac120012\tfrac12
dydx\dfrac{dy}{dx}++00++

The gradient does not change sign, so (0,−6)(0, -6) is a point of inflexion.

Example 4.20.

Sketch the curve y=2x3+x2−4x+1y = 2x^3 + x^2 - 4x + 1.

Finding the maximum and minimum turning points gives a general idea of the shape and position of the curve.

y=2x3+x2−4x+1,dydx=6x2+2x−4,d2ydx2=12x+2.y = 2x^3 + x^2 - 4x + 1, \qquad \frac{dy}{dx} = 6x^2 + 2x - 4, \qquad \frac{d^2 y}{dx^2} = 12x + 2 .

At turning points 6x2+2x−4=06x^2 + 2x - 4 = 0, that is 2(3x−2)(x+1)=02(3x - 2)(x + 1) = 0, so x=23x = \tfrac23 or x=−1x = -1.

  1. When x=23x = \tfrac23, d2ydx2=12(23)+2>0\dfrac{d^2 y}{dx^2} = 12\left(\tfrac23\right) + 2 > 0 and y=2(23)3+(23)2−4(23)+1=−1727y = 2\left(\tfrac23\right)^3 + \left(\tfrac23\right)^2 - 4\left(\tfrac23\right) + 1 = -\tfrac{17}{27}, so (23,−1727)\left(\tfrac23, -\tfrac{17}{27}\right) is a minimum turning point.
  2. When x=−1x = -1, d2ydx2=12(−1)+2<0\dfrac{d^2 y}{dx^2} = 12(-1) + 2 < 0 and y=2(−1)3+(−1)2−4(−1)+1=4y = 2(-1)^3 + (-1)^2 - 4(-1) + 1 = 4, so (−1,4)(-1, 4) is a maximum turning point.
  3. As y=2x3+x2−4x+1y = 2x^3 + x^2 - 4x + 1, the curve cuts the yy-axis at (0,1)(0, 1).

From these three results we can sketch the curve.

−2−11−2246xy(−1, 4)(⅔, −17/27)(0, 1)
Figure 4.12. A sketch of y=2x3+x2−4x+1y = 2x^3 + x^2 - 4x + 1 from its turning points and its intercept on the yy-axis.

In some cases the intercepts on the xx-axis can be found as well, although the equation which gives them is not always easy to solve. Here it is 2x3+x2−4x+1=02x^3 + x^2 - 4x + 1 = 0.

Example 4.21.

A farmer has an adjustable electric fence that is 100100 m long. It is used to enclose a rectangular grazing area on three sides, the fourth side being a fixed hedge. Find the maximum area that can be enclosed.

The length of the enclosure can be varied, up to 100100 m, and the width and area then depend on the length chosen. Let the length SRSR be xx m. Then as PS+SR+RQ=100PS + SR + RQ = 100,

PS=RQ=12(100−x),PS = RQ = \tfrac12(100 - x),

and the area AA of the enclosure, in square metres, is

A=x[12(100−x)]=50x−12x2.A = x\left[\tfrac12(100 - x)\right] = 50x - \tfrac12 x^2 .

Now AA is a quadratic function of xx with a negative coefficient of x2x^2, so from what we know of quadratic functions it is greatest at x=50x = 50, halfway between its zeros at x=0x = 0 and x=100x = 100. So the maximum area that can be enclosed is 50×25=125050 \times 25 = 1250 square metres.

The maximum value of 50x−12x250x - \tfrac12 x^2 can also be found by differentiation. We have

dAdx=50−x,\frac{dA}{dx} = 50 - x,

and AA has a stationary value when dAdx=0\dfrac{dA}{dx} = 0, that is when x=50x = 50. Now d2Adx2=−1\dfrac{d^2 A}{dx^2} = -1, which is negative, so this is a maximum, and the maximum value is 25×50=125025 \times 50 = 1250 square metres.

hedgePQSRx½(100 − x)½(100 − x)the enclosure501005001000xA(50, 1250)its area
Figure 4.13. The enclosure against the hedge, and its area A=50x−12x2A = 50x - \tfrac12 x^2 as the length xx varies.

The second method, using differentiation, is necessary when finding maximum or minimum values of functions that are not quadratic. For quadratic functions the first method is preferable, since their properties let us find the maximum or minimum value by inspection. In this case the greatest value and the maximum value are the same.

Example 4.22.

A rectangle has a perimeter of 2020 units. Find its greatest possible area.

The perimeter, the total length of the edges, is twice the short side plus twice the long side, so the two sides add up to 1010. So if one side has length xx, the other has length 10−x10 - x, and the area of the rectangle is

x(10−x)=10x−x2.x(10 - x) = 10x - x^2 .

This is differentiable everywhere, with derivative 10−2x10 - 2x. At its greatest value, either its gradient is 00 or xx is at one of the ends of its range, x=0x = 0 or x=10x = 10. At the ends the area is 00, so we solve 10−2x=010 - 2x = 0, which gives x=5x = 5. The area there is 10(5)−52=2510(5) - 5^2 = 25, so the greatest area is 2525 square units.

The point where the gradient is zero is a maximum and not a minimum, for three separate reasons.

  1. x=5x = 5 is the only place where 10x−x210x - x^2 can turn round, and at values of xx on either side of it, such as 00 and 1010, the area is less than 2525. If the area were more than 2525 at x=6x = 6, say, the curve would have to turn round again to reach (10,0)(10, 0), and it does not.
  2. The gradient 10−2x10 - 2x is positive to the left of x=5x = 5 and negative to the right.
  3. The graph of 10x−x210x - x^2 is a parabola opening downwards, so its only flat point is a maximum.

Remark.

If f(x)=ax2+bx+cf(x) = ax^2 + bx + c, then f′(x)=2ax+bf'(x) = 2ax + b, so f(x)f(x) has a stationary value where x=−b2ax = -\dfrac{b}{2a}, on the axis of the curve. If α\alpha and β\beta are the roots of ax2+bx+c=0ax^2 + bx + c = 0, where the curve crosses the xx-axis, then by the sum of the roots 12(α+β)=−b2a\tfrac12(\alpha + \beta) = -\dfrac{b}{2a}, so the turning point of the curve has xx-coordinate 12(α+β)\tfrac12(\alpha + \beta), halfway between the roots. A ball thrown through the air follows such a parabola, if we ignore air resistance.

Problem 4.8.

Find the stationary values of each of the following functions and determine their nature.

  1. 2x3−3x2−12x2x^3 - 3x^2 - 12x;
  2. (x−3)(2x+1)(x - 3)(2x + 1);
  3. x3+3x^3 + 3.

The Number ee

Exponential Growth

The curve y=2xy = 2^x of the last lesson increases as xx increases. Its values at x=−3,−2,−1,0,1,2,3x = -3, -2, -1, 0, 1, 2, 3 are 18,14,12,1,2,4,8\tfrac18, \tfrac14, \tfrac12, 1, 2, 4, 8, and the graph grows quite fast. Differentiating from first principles shows how fast:

2x+h−2xh=2x⋅2h−2xh=2x⋅2h−1h,\frac{2^{x + h} - 2^x}{h} = \frac{2^x \cdot 2^h - 2^x}{h} = 2^x \cdot \frac{2^h - 1}{h},

and the factor 2h−1h\dfrac{2^h - 1}{h} does not involve xx at all. Its values, and those of the corresponding factor for 3x3^x, as hh approaches zero are

hh0.10.10.010.010.0010.0010.00010.0001
2h−1h\dfrac{2^h - 1}{h}0.71770.71770.69560.69560.69340.69340.69320.6932
3h−1h\dfrac{3^h - 1}{h}1.16121.16121.10471.10471.09921.09921.09871.0987

so the gradient of 2x2^x is about 0.693×2x0.693 \times 2^x, and the gradient of 3x3^x is about 1.099×3x1.099 \times 3^x.

The rate of change of an exponential function is proportional to its current value. The larger the function gets, the faster it grows, which makes it larger still; this runaway growth is called exponential growth. It is what happens with the spread of a disease, or with interest in a bank.

A base for which the factor is exactly 11 would give a function whose derivative is equal to the function itself. The factor is less than 11 for base 22 and greater than 11 for base 33, and there is a number between 22 and 33 for which it is exactly 11.

Definition 4.23 (The Number e).

The number ee, which is about 2.71828182845904523536…2.71828182845904523536\ldots, is the base for which

ddx(ex)=ex.\frac{d}{dx}\left(e^x\right) = e^x .

With ee in place of 22 the factor takes the values 1.05171.0517, 1.00501.0050, 1.00051.0005 and 1.00011.0001 for the four values of hh in the table. That such a number exists, and that exe^x has exactly this derivative, we accept on trust at this stage, as we did the power rule for fractional powers.

−2−1121234567xygradient 1gradient ey = ex
Figure 4.14. The curve y=exy = e^x. At every point the gradient equals the height: the tangent at (0,1)(0, 1) has gradient 11, and the tangent at (1,e)(1, e) has gradient ee.

Example 4.24.

Find the equation of the tangent to y=exy = e^x at the point (0,1)(0, 1).

Since dydx=ex\dfrac{dy}{dx} = e^x, the gradient at (0,1)(0, 1) is e0=1e^0 = 1. The tangent is the line of gradient 11 through (0,1)(0, 1), which is y=x+1y = x + 1.

The Natural Logarithm

Definition 4.25 (Natural Logarithm).

The natural logarithm of a positive number xx is its logarithm to base ee, written

ln⁡x=log⁡ex.\ln x = \log_e x .

So ln⁡x\ln x is the number yy for which ey=xe^y = x. Since log⁡a1=0\log_a 1 = 0 and log⁡a(ak)=k\log_a\left(a^k\right) = k for every base, straight from the definition of a logarithm, we have ln⁡1=0\ln 1 = 0, ln⁡e=1\ln e = 1 and ln⁡(ek)=k\ln\left(e^k\right) = k.

As a function from the positive numbers to the real numbers, ln⁡x\ln x is the inverse of the function exe^x from the real numbers to the positive numbers, in the same way that log⁡2x\log_2 x is the inverse of 2x2^x. Its graph is the reflection of y=exy = e^x in the line y=xy = x, and has the same shape as the graph of y=log⁡2xy = \log_2 x.

Remark.

People often write log⁡x\log x with no base. If you are a physicist or an engineer this usually means log⁡10x\log_{10} x, but mathematicians usually mean ln⁡x\ln x. In these notes base 1010 is written lg⁡\lg, base ee is written ln⁡\ln, and log⁡\log without a base appears only in rules that hold for every base.

Example 4.26.

Find the coordinates of the stationary point of y=ex−3xy = e^x - 3x, and determine its nature.

dydx=ex−3,d2ydx2=ex.\frac{dy}{dx} = e^x - 3, \qquad \frac{d^2 y}{dx^2} = e^x .

The gradient is zero when ex=3e^x = 3, that is when x=ln⁡3x = \ln 3, and there y=eln⁡3−3ln⁡3=3−3ln⁡3y = e^{\ln 3} - 3\ln 3 = 3 - 3\ln 3. Since exe^x is positive for every xx, d2ydx2>0\dfrac{d^2 y}{dx^2} > 0, so (ln⁡3,3−3ln⁡3)(\ln 3, 3 - 3\ln 3), which is about (1.099,−0.296)(1.099, -0.296), is a minimum turning point.

Problem 4.9.

  1. Solve the equation ex=10e^x = 10, giving xx exactly and to three decimal places.
  2. Differentiate x3+exx^3 + e^x with respect to xx.
  3. Find the gradient of the curve y=ex−x2y = e^x - x^2 at the point where x=0x = 0.

Exercises

Exercise 4.1.

Differentiate f(x)=x+1xf(x) = x + \dfrac1x from first principles.

Exercise 4.2.

Find the equations of the tangents to the curve y=(2x−1)(x+1)y = (2x - 1)(x + 1) at the points where the curve cuts the xx-axis, and find the point of intersection of these tangents.

Exercise 4.3.

Find the coordinates of the point on y=x2−5y = x^2 - 5 at which the gradient is 33. Hence find the value of cc for which the line y=3x+cy = 3x + c is a tangent to y=x2−5y = x^2 - 5.

Exercise 4.4.

Find the value of kk for which y=2x+ky = 2x + k is a normal to y=2x2−3y = 2x^2 - 3.

Exercise 4.5.

Find the coordinates of the point on y=x2−7x+3y = x^2 - 7x + 3 at which the gradient is 22. Hence find the equation of the normal to y=x2−7x+3y = x^2 - 7x + 3 which is parallel to x+2y−1=0x + 2y - 1 = 0.

Exercise 4.6.

Show that (t,1t)\left(t, \dfrac1t\right) lies on the curve y=1xy = \dfrac1x for all non-zero values of tt. Find the equation of the tangent to y=1xy = \dfrac1x at (t,1t)\left(t, \dfrac1t\right), and the area of the triangle enclosed by this tangent and the coordinate axes.

Exercise 4.7.

Find the turning points on y=3x4+4x3−12x2y = 3x^4 + 4x^3 - 12x^2, and give a rough sketch of the curve.

Exercise 4.8.

Find the stationary values of f(x)=x+1xf(x) = x + \dfrac1x, and use them to sketch the curve y=x+1xy = x + \dfrac1x.

Exercise 4.9.

Find the coordinates of the turning points of the curve y=3+24x−21x2−4x3y = 3 + 24x - 21x^2 - 4x^3, and sketch the curve.

Exercise 4.10.

The gradient function of y=ax2+bx+cy = ax^2 + bx + c is 4x+24x + 2, and the function has a minimum value of 11. Find the values of aa, bb and cc.

Exercise 4.11.

A closed cylindrical can has height hh and base radius rr, and its volume is 0.010.01 m³. Show that

h=1100πr2.h = \frac{1}{100\pi r^2} .

Show further that SS, the surface area, is given by

S=2πr2+150r,S = 2\pi r^2 + \frac{1}{50r},

and hence find the value of rr for which SS is a minimum.

Exercise 4.12.

An open rectangular box is made from a square sheet of cardboard by removing a square from each corner and joining the cut edges. If the cardboard has edge 0.50.5 m, find the maximum volume of the box.

Exercise 4.13.

A cylinder is cut from a solid sphere of radius 55 cm, so that the curved edges of the cylinder reach the surface of the sphere. If the height of the cylinder is 2h2h, show that its volume is 2πh(25−h2)2\pi h\left(25 - h^2\right), and find the maximum volume of such a cylinder.

Exercise 4.14.

A piece of string of fixed length is made to enclose a rectangle. Show that the enclosed area is greatest when the rectangle is a square.

Exercise 4.15.

The table shows the results of an experiment in which some hot liquid was left to cool, its temperature θ\theta °C being measured at intervals of one minute.

Time tt (minutes)001122334455
Temperature θ\theta (°C)10010093938787838380807878

What does dθdt\dfrac{d\theta}{dt} represent? By drawing a graph of these results, estimate the rate of decrease of the temperature after three minutes.

Exercise 4.16.

Show that the tangent to y=exy = e^x at the point where x=1x = 1 passes through the origin. Hence find the value of kk for which the line y=kxy = kx is a tangent to y=exy = e^x.

Check Yourself

 

Fresh questions on the whole chapter — none of them is worked out above. Do each on paper first; the box only tells you whether you got there.

Answers are checked in your browser, as often as you like. Nothing is sent anywhere and nothing is kept but your own work. A formula may be written with the symbols themselves or with ~ & | -> <-> ^, and \and, \or, \to expand as you type.

Exercise 4.17.

What is lim⁡x→3x2−9x−3\displaystyle\lim_{x \to 3} \frac{x^2 - 9}{x - 3}?

answer one of these

Exercise 4.18.

What is lim⁡x→∞3x2\displaystyle\lim_{x \to \infty} \frac{3}{x^2}?

answer one of these

Exercise 4.19.

What is the gradient of the chord of y=x2y = x^2 joining the points where x=1x = 1 and x=3x = 3?

answer one of these

Exercise 4.20.

What is ddx(x5)\dfrac{d}{dx}\left(x^5\right)?

answer one of these

Exercise 4.21.

What is ddx(x5/2)\dfrac{d}{dx}\left(x^{5/2}\right)?

answer one of these

Exercise 4.22.

What is ddx(4x)\dfrac{d}{dx}\left(\dfrac4x\right)?

answer one of these

Exercise 4.23.

What is ddx(x+2)2\dfrac{d}{dx}(x + 2)^2?

answer one of these

Exercise 4.24.

What is the gradient of the curve y=x3−xy = x^3 - x at the point where x=2x = 2?

answer one of these

Exercise 4.25.

What is the equation of the tangent to y=x2y = x^2 at the point (2,4)(2, 4)?

answer one of these

Exercise 4.26.

What is the gradient of the normal to y=x2y = x^2 at the point where x=1x = 1?

answer one of these

Exercise 4.27.

At which values of xx does x3−3xx^3 - 3x have stationary values?

answer one of these

Exercise 4.28.

At the origin, the curve y=x4y = x^4 has

answer one of these

Exercise 4.29.

A rectangle has a perimeter of 1212 units. What is its greatest possible area?

answer one of these

Exercise 4.30.

For which values of xx is x2−4xx^2 - 4x decreasing?

answer one of these

Exercise 4.31.

What is ddx(ex+x2)\dfrac{d}{dx}\left(e^x + x^2\right)?

answer one of these

Exercise 4.32.

What is the solution of e2x=7e^{2x} = 7?

answer one of these

Lesson 5

Trigonometric Functions, and Identities

Taught

Measuring Angles

Measurement of Rotation

When a line OPOP is pivoted at OO and rotates from its initial position OP0OP_0 to a new position OP1OP_1, the angle P0OP1P_0OP_1 is a measure of the rotation of OPOP.

OP0P1θ
Figure 5.1. The angle θ\theta measures the rotation of OPOP from OP0OP_0 to OP1OP_1.

The angle θ\theta is usually measured in one of two units.

The Degree

The ancient Babylonian mathematicians, who thought that the solar year was 360360 days long, divided one complete revolution into 360360 equal parts, each part now being known as one degree, 1∘1^\circ. Using the degree as the unit of rotation, half a revolution corresponds to 180∘180^\circ and a quarter of a revolution, that is a right angle, corresponds to 90∘90^\circ.

Angles smaller than a degree are usually given as decimal parts, so that half a degree is 0.5∘0.5^\circ. In some fields, such as navigation, a degree is divided into 6060 minutes, written 60′60', and each minute into 6060 seconds, written 60′′60''.

The Radian

Definition 5.1 (Radian).

If an arc PP1PP_1 of a circle with centre OO is equal in length to the radius of the circle, the angle POP1POP_1 is one radian, written 1c1^{\mathrm{c}}.

OPP1rrr1 rad
Figure 5.2. An arc equal in length to the radius subtends one radian at the centre.

The number of radians in one complete revolution is therefore

circumferenceradius=2πrr=2π,\frac{\text{circumference}}{\text{radius}} = \frac{2\pi r}{r} = 2\pi,

since the circumference of a circle of radius rr is 2πr2\pi r. So

2π radians=360∘,π radians=180∘,π2 radians=90∘.2\pi \text{ radians} = 360^\circ, \qquad \pi \text{ radians} = 180^\circ, \qquad \frac{\pi}{2} \text{ radians} = 90^\circ .

When an angle is quoted in terms of π\pi it is normal to omit the radian symbol, so we write 180∘=π180^\circ = \pi, not πc\pi^{\mathrm{c}}.

Angles which are simple fractions of 180∘180^\circ are easily expressed in radians in terms of π\pi, using 180∘=π180^\circ = \pi, and conversely:

60∘=60×π180=π3,135∘=135×π180=3π4,7π6=7π6×180∘π=210∘,5π3=5π3×180∘π=300∘.\begin{aligned} 60^\circ &= 60 \times \frac{\pi}{180} = \frac{\pi}{3}, & 135^\circ &= 135 \times \frac{\pi}{180} = \frac{3\pi}{4}, \\ \frac{7\pi}{6} &= \frac{7\pi}{6} \times \frac{180^\circ}{\pi} = 210^\circ, & \frac{5\pi}{3} &= \frac{5\pi}{3} \times \frac{180^\circ}{\pi} = 300^\circ . \end{aligned}

Converting in this way is not convenient for angles which are not simple fractions of a revolution. It would not be easy, for instance, to express 47∘34′47^\circ 34' as a multiple of π\pi. We would first have to express 34′34' as a decimal part of a degree, 34′=0.567∘34' = 0.567^\circ, and then use

47.567∘=47.567×π180=0.830 radians.47.567^\circ = 47.567 \times \frac{\pi}{180} = 0.830 \text{ radians}.

Calculations like this are tedious, and a calculator which offers the conversion avoids them.

To visualise the size of an angle of one radian, it helps to remember that π\pi radians =180∘= 180^\circ and that π=3.14\pi = 3.14 to two decimal places, so

1 radian=180∘3.14=57.3∘to three significant figures,1 \text{ radian} = \frac{180^\circ}{3.14} = 57.3^\circ \quad\text{to three significant figures},

a little less than 60∘60^\circ.

Problem 5.1.

  1. Without a calculator, express 30∘30^\circ, 270∘270^\circ and 22.5∘22.5^\circ in radians, in terms of π\pi.
  2. Without a calculator, express 5π6\dfrac{5\pi}{6}, 11π6\dfrac{11\pi}{6} and 4π9\dfrac{4\pi}{9} in degrees.
  3. Express 54∘45′54^\circ 45' in radians, and 2.862.86 radians in degrees and minutes.

The Circular Functions

The Ratios of an Acute Angle

For any acute angle θ\theta there are six trigonometric ratios, each of which is defined by referring to a right-angled triangle containing θ\theta.

OABθ
Figure 5.3. A right-angled triangle OABOAB containing the acute angle θ\theta at OO.

Definition 5.2 (Trigonometric Ratios of an Acute Angle).

In a triangle OABOAB with a right angle at AA and an acute angle θ\theta at OO,

sin⁡θ=ABOB,cosec⁡θ=OBAB,cos⁡θ=OAOB,sec⁡θ=OBOA,tan⁡θ=ABOA,cot⁡θ=OAAB.\begin{aligned} \sin\theta &= \frac{AB}{OB}, &\quad \operatorname{cosec}\theta &= \frac{OB}{AB}, \\ \cos\theta &= \frac{OA}{OB}, &\quad \sec\theta &= \frac{OB}{OA}, \\ \tan\theta &= \frac{AB}{OA}, &\quad \cot\theta &= \frac{OA}{AB} . \end{aligned}

These are the sine, cosine, tangent, cosecant, secant and cotangent of θ\theta.

Each ratio has a single value for any one acute angle, since all right-angled triangles containing θ\theta are similar, and the values are available from a calculator.

Length of an Arc and Area of a Sector

The circumference of a circle of radius rr is 2πr2\pi r and its area is πr2\pi r^2, and these formulae can be used to derive further results.

Consider an arc which subtends an angle θ\theta at the centre of the circle, where θ\theta is measured in radians. From the definition of a radian, the arc which subtends 11 radian at the centre has length rr, so an arc which subtends θ\theta radians at the centre has length rθr\theta.

The area of a sector containing an angle of θ\theta radians at the centre can be found by regarding the sector as a fraction of the circle. The ratio of the area of the sector to the area of the circle is equal to the ratio of the angle θ\theta contained in the sector to the angle 2π2\pi contained in the whole circle, that is

area of sectorπr2=θ2π,soarea of sector=12r2θ.\frac{\text{area of sector}}{\pi r^2} = \frac{\theta}{2\pi}, \qquad\text{so}\qquad \text{area of sector} = \tfrac12 r^2\theta .
ABOθrrθ
Figure 5.4. A sector AOBAOB containing the angle θ\theta at the centre OO.

Hence if an arc ABAB subtends an angle of θ\theta radians at the centre OO of a circle of radius rr,

length of arc AB=rθ,area of sector AOB=12r2θ.\text{length of arc } AB = r\theta, \qquad \text{area of sector } AOB = \tfrac12 r^2\theta .

Example 5.3.

A railway line changes direction by 20∘20^\circ when passing round a circular arc of length 500500 m. What is the radius of the arc?

The line turns through the angle which the arc subtends at the centre, and in radians 20∘=20×π180=π920^\circ = 20 \times \dfrac{\pi}{180} = \dfrac{\pi}{9}. Using length of arc =rθ= r\theta, with θ\theta in radians,

500=r×π9,sor=4500π=1432.500 = r \times \frac{\pi}{9}, \qquad\text{so}\qquad r = \frac{4500}{\pi} = 1432 .

The radius of the track is therefore 14321432 m.

Example 5.4.

A chord ABAB divides a circle of radius 22 m into two segments. If ABAB subtends an angle of 60∘60^\circ at the centre OO of the circle, find the area of the minor segment.

The triangle AOBAOB has base OA=2OA = 2 and height 2sin⁡60∘2\sin 60^\circ, so

area of triangle AOB=12×2×2sin⁡60∘=1.732 m2.\text{area of triangle } AOB = \tfrac12 \times 2 \times 2\sin 60^\circ = 1.732\ \text{m}^2 .

Since 60∘=π360^\circ = \dfrac{\pi}{3} radians,

area of sector AOB=12×22×π3=2π3=2.094 m2.\text{area of sector } AOB = \tfrac12 \times 2^2 \times \frac{\pi}{3} = \frac{2\pi}{3} = 2.094\ \text{m}^2 .

So the area of the minor segment, shaded in Figure 5.5, is (2.094−1.732) m2=0.362 m2(2.094 - 1.732)\ \text{m}^2 = 0.362\ \text{m}^2.

ABO60°2 m
Figure 5.5. The minor segment cut off by a chord subtending 60∘60^\circ at the centre.

Example 5.5.

Two discs of radii 33 cm and 44 cm are laid on a table with their centres 55 cm apart. Find the perimeter of the figure-eight shape so formed.

Let the discs have their centres at AA and BB and cross at CC. Since 52=32+425^2 = 3^2 + 4^2, the triangle ABCABC is right-angled at CC. Let α\alpha be the angle CABCAB and β\beta the angle CBACBA. Then

tan⁡α=43,α=53.13∘=0.927 radians,β=90∘−α=36.87∘=0.644 radians.\tan\alpha = \tfrac43, \quad \alpha = 53.13^\circ = 0.927 \text{ radians}, \qquad \beta = 90^\circ - \alpha = 36.87^\circ = 0.644 \text{ radians}.

The perimeter is made up of an arc subtending 2π−2α2\pi - 2\alpha radians in the circle of radius 33 cm and an arc subtending 2π−2β2\pi - 2\beta radians in the circle of radius 44 cm. Hence, using length of arc =rθ= r\theta,

perimeter=3(6.283−1.854)+4(6.283−1.287)=33.3 cm.\text{perimeter} = 3(6.283 - 1.854) + 4(6.283 - 1.287) = 33.3 \text{ cm}.
ABCαβ345
Figure 5.6. The figure-eight formed by the two discs, with its perimeter drawn heavily.

Problem 5.2.

  1. The moon subtends an angle of 31′31' at the earth, and its distance from the earth is 382 100382\,100 km. Find the diameter of the moon in kilometres.
  2. An arc ABAB of length 55 cm is marked on a circle of radius 33 cm. Find the area of the sector bounded by this arc and the radii to AA and BB.
  3. A chord ABAB of length 5.25.2 cm subtends an angle of 120∘120^\circ at the centre of a circle. Calculate the length of the arc ABAB, the area of the sector containing the angle of 120∘120^\circ, and the area of the minor segment cut off by ABAB.

The Ratios of a General Angle

Since we now regard an angle as the measure of the rotation of a line about a fixed point, the size of an angle is unlimited, because the line can keep on rotating indefinitely. The six trigonometric ratios, however, have so far been given a meaning only for acute angles, since each is defined by an angle in a right-angled triangle. To use them for angles of any size they must be defined in a more general way.

The system of reference in which a general angle is measured is very similar to that used for polar coordinates. The point about which the line OPOP rotates is the pole or origin OO, and the position from which the angle is measured is the initial line, the xx-axis. An angle formed when the line rotates anticlockwise is positive, while clockwise rotation gives a negative angle. The pair of Cartesian axes divides the plane into four quadrants, numbered 11, 22, 33 and 44 as in Figure 5.7.

xy1234Pθanticlockwise: θ positivexyPθclockwise: θ negative
Figure 5.7. The four quadrants, and the positive and negative senses of rotation.

As the line OPOP rotates, PP moves round the first quadrant, where both its coordinates are positive. As OPOP moves into the second quadrant its xx-coordinate becomes negative. In the third quadrant both coordinates of PP are negative, and in the fourth quadrant the xx-coordinate is positive and the yy-coordinate negative. The length rr of OPOP, the radius vector, is always taken to be positive.

Definition 5.6 (Trigonometric Ratios of Any Angle).

If the line OPOP has rotated through the angle θ\theta from the positive xx-axis, and PP is the point (x,y)(x, y) with OP=rOP = r, then

sin⁡θ=yr,cos⁡θ=xr,tan⁡θ=yx,\sin\theta = \frac{y}{r}, \qquad \cos\theta = \frac{x}{r}, \qquad \tan\theta = \frac{y}{x},

and cosec⁡θ=ry\operatorname{cosec}\theta = \dfrac{r}{y}, sec⁡θ=rx\sec\theta = \dfrac{r}{x}, cot⁡θ=xy\cot\theta = \dfrac{x}{y}, each ratio being defined whenever its denominator is not zero.

For an acute angle these agree with the ratios of the right-angled triangle formed by OO, PP and the foot of the perpendicular from PP to the xx-axis.

The numerical values of the ratios can be found as follows. From any position of PP, the perpendicular from PP meeting the xx-axis at QQ forms a right-angled triangle OPQOPQ. The angle POQPOQ formed in this way is always acute, whatever the value of θ\theta, and is called the associated acute angle, α\alpha. For a particular value of θ\theta, α\alpha is the difference between θ\theta and 180∘180^\circ or 360∘360^\circ, or a further multiple of 180∘180^\circ for larger angles. For example

θ=160∘:α=180∘−160∘=20∘,θ=7π6:α=7π6−π=π6,θ=275∘:α=360∘−275∘=85∘,θ=517∘:α=540∘−517∘=23∘.\begin{aligned} \theta &= 160^\circ: & \alpha &= 180^\circ - 160^\circ = 20^\circ, &\qquad \theta &= \tfrac{7\pi}{6}: & \alpha &= \tfrac{7\pi}{6} - \pi = \tfrac{\pi}{6}, \\ \theta &= 275^\circ: & \alpha &= 360^\circ - 275^\circ = 85^\circ, &\qquad \theta &= 517^\circ: & \alpha &= 540^\circ - 517^\circ = 23^\circ . \end{aligned}
xyαP(x, y)Q1st quadrant: α = θxyαP(x, y)Q2nd quadrant: α = 180° − θxyαP(x, y)Q3rd quadrant: α = θ − 180°xyαP(x, y)Q4th quadrant: α = 360° − θ
Figure 5.8. The associated acute angle α\alpha when PP is in each of the four quadrants.

In the triangle OPQOPQ the lengths PQPQ and OQOQ are rsin⁡αr\sin\alpha and rcos⁡αr\cos\alpha, so the ratios of θ\theta have the numerical values of the ratios of α\alpha, and their signs depend on the signs of xx and yy, that is on the quadrant into which PP has rotated.

  1. In the first quadrant all six ratios are positive, and since θ\theta is acute their values are those of θ\theta itself. If OPOP has rotated through more than a complete revolution, α=θ−360∘\alpha = \theta - 360^\circ.
  2. In the second quadrant α=180∘−θ\alpha = 180^\circ - \theta. The sine ratio yr\dfrac{y}{r} is positive, the cosine ratio xr\dfrac{x}{r} is negative and the tangent ratio yx\dfrac{y}{x} is negative, so sin⁡θ=+sin⁡α\sin\theta = +\sin\alpha, cos⁡θ=−cos⁡α\cos\theta = -\cos\alpha and tan⁡θ=−tan⁡α\tan\theta = -\tan\alpha.
  3. In the third quadrant α=θ−180∘\alpha = \theta - 180^\circ. The tangent ratio is positive while the sine and cosine ratios are both negative, so sin⁡θ=−sin⁡α\sin\theta = -\sin\alpha, cos⁡θ=−cos⁡α\cos\theta = -\cos\alpha and tan⁡θ=+tan⁡α\tan\theta = +\tan\alpha.
  4. In the fourth quadrant α=360∘−θ\alpha = 360^\circ - \theta. The cosine ratio is positive but the sine and tangent ratios are negative, so sin⁡θ=−sin⁡α\sin\theta = -\sin\alpha, cos⁡θ=+cos⁡α\cos\theta = +\cos\alpha and tan⁡θ=−tan⁡α\tan\theta = -\tan\alpha.

These results are summarised in the quadrant diagrams of Figure 5.9. The second shows in each quadrant the ratios which are positive there: all of them in the first, the sine in the second, the tangent in the third and the cosine in the fourth. The quadrant rule, together with the value of the associated acute angle, gives the value of any trigonometric ratio of any angle.

s +c +t +s +c −t −s −c −t +s −c +t −(i) the signsASTC(ii) the ratios that are positive
Figure 5.9. Quadrant diagrams for the signs of the sine, cosine and tangent.

Negative Angles

If OPOP rotates clockwise, so that PP moves through the quadrants in the reverse order, 4th, 3rd, 2nd, 1st, then θ\theta is negative. Every position of OPOP can be reached either by anticlockwise or by clockwise rotation, and so corresponds to two different values of θ\theta, one positive and one negative. For example, θ=+120∘\theta = +120^\circ and θ=−240∘\theta = -240^\circ give the same position, with α=60∘\alpha = 60^\circ in both cases, so the angles +120∘+120^\circ and −240∘-240^\circ have the same trigonometric ratios. PP is in the second quadrant, where only the sine ratio is positive, hence

sin⁡120∘=sin⁡(−240∘)=+sin⁡60∘,cos⁡120∘=cos⁡(−240∘)=−cos⁡60∘,tan⁡120∘=tan⁡(−240∘)=−tan⁡60∘.\begin{aligned} \sin 120^\circ &= \sin(-240^\circ) = +\sin 60^\circ, \\ \cos 120^\circ &= \cos(-240^\circ) = -\cos 60^\circ, \\ \tan 120^\circ &= \tan(-240^\circ) = -\tan 60^\circ . \end{aligned}
xyP+120°−240°
Figure 5.10. The same position of OPOP reached by rotations of +120∘+120^\circ and −240∘-240^\circ.

Turning through −θ-\theta instead of θ\theta reflects PP in the xx-axis, which changes the sign of yy and leaves xx unchanged. So for every angle

sin⁡(−θ)=−sin⁡θ,cos⁡(−θ)=cos⁡θ,tan⁡(−θ)=−tan⁡θ.\sin(-\theta) = -\sin\theta, \qquad \cos(-\theta) = \cos\theta, \qquad \tan(-\theta) = -\tan\theta .

Example 5.7.

Find the sine, cosine and tangent of 243∘243^\circ.

Here θ=243∘\theta = 243^\circ and α=θ−180∘=63∘\alpha = \theta - 180^\circ = 63^\circ. PP is in the third quadrant, where only the tangent ratio is positive, so

sin⁡243∘=−sin⁡63∘=−0.8910,cos⁡243∘=−cos⁡63∘=−0.4540,tan⁡243∘=+tan⁡63∘=1.9626.\sin 243^\circ = -\sin 63^\circ = -0.8910, \qquad \cos 243^\circ = -\cos 63^\circ = -0.4540, \qquad \tan 243^\circ = +\tan 63^\circ = 1.9626 .

Example 5.8.

If cos⁡θ=0.866\cos\theta = 0.866 and tan⁡θ\tan\theta is negative, find sin⁡θ\sin\theta.

θ\theta is in a quadrant where the cosine ratio is positive and the tangent ratio is negative, that is in the fourth quadrant, where α=360∘−θ\alpha = 360^\circ - \theta. Now cos⁡α=0.866\cos\alpha = 0.866 gives α=30∘\alpha = 30^\circ, so θ=330∘\theta = 330^\circ. In the fourth quadrant the sine ratio is negative, therefore

sin⁡θ=−sin⁡α=−sin⁡30∘=−0.5.\sin\theta = -\sin\alpha = -\sin 30^\circ = -0.5 .

Example 5.9.

Given that tan⁡θ=1\tan\theta = 1 and −2π⩽θ⩽2π-2\pi \leqslant \theta \leqslant 2\pi, give four possible values of θ\theta.

The tangent ratio is positive in the first and third quadrants. When tan⁡α=1\tan\alpha = 1, α=π4\alpha = \dfrac{\pi}{4}, or 45∘45^\circ. The range of values of θ\theta is specified in radians, so the solution should also be given in radians:

θ=π4,5π4,−3π4or−7π4.\theta = \frac{\pi}{4}, \qquad \frac{5\pi}{4}, \qquad -\frac{3\pi}{4} \qquad\text{or}\qquad -\frac{7\pi}{4} .

In this example we are solving a simple trigonometric equation.

Example 5.10.

An angle θ\theta has an associated acute angle of 53∘53^\circ. Find sin⁡θ\sin\theta.

Only α\alpha is given, so θ\theta could be in any of the four quadrants. In quadrants 11 and 22, sin⁡θ=sin⁡53∘=0.7986\sin\theta = \sin 53^\circ = 0.7986, and in quadrants 33 and 44, sin⁡θ=−sin⁡53∘=−0.7986\sin\theta = -\sin 53^\circ = -0.7986. Therefore sin⁡θ=±0.7986\sin\theta = \pm 0.7986.

Problem 5.3.

  1. Find the sine, cosine and tangent of 300∘300^\circ and of −160∘-160^\circ in terms of the ratios of their associated acute angles, and evaluate them.
  2. Within the range −360∘⩽θ⩽360∘-360^\circ \leqslant \theta \leqslant 360^\circ, give all the values of θ\theta for which cos⁡θ=−0.5\cos\theta = -0.5, and all those for which tan⁡θ=1.2\tan\theta = 1.2.
  3. Find the smallest angle, positive or negative, for which cos⁡θ=0.8\cos\theta = 0.8 and sin⁡θ\sin\theta is positive, and the smallest for which sin⁡θ=−0.6\sin\theta = -0.6 and tan⁡θ\tan\theta is negative.

Solving Triangles

Triangles are involved in many practical measurements, in surveying for instance. A triangle has three sides and three angles, and its size and shape can be specified by suitable data, such as three sides, two angles and a side, or two sides and their included angle. The third angle is not an independent item: the angles of a triangle add up to 180∘180^\circ, so if two of them are known the third follows directly. Two sides and a non-included angle can sometimes give two different triangles. From any data sufficient to define a triangle the remaining sides and angles can be calculated. This is called solving the triangle, and it uses one of the formulae that relate the sides and angles of a triangle. The two used most frequently are the sine rule and the cosine rule.

In a triangle ABCABC we use AA, BB and CC to denote the angles at the vertices AA, BB and CC, and aa, bb and cc to denote the sides opposite these vertices.

The Sine Rule

In a triangle ABCABC,

asin⁡A=bsin⁡B=csin⁡C.\frac{a}{\sin A} = \frac{b}{\sin B} = \frac{c}{\sin C} .

We show this by taking OO and RR as the centre and the radius of the circle through AA, BB and CC, drawing the diameter ADAD and joining DBDB, as in Figure 5.11. Then ∠ABD=90∘\angle ABD = 90^\circ, the angle in a semicircle, so in the right-angled triangle ABDABD, whose hypotenuse ADAD is 2R2R,

c=2Rsin⁡D.c = 2R\sin D .

In diagram (i), C=DC = D, since they are angles in the same segment. In diagram (ii), C=180∘−DC = 180^\circ - D, since they are opposite angles of a cyclic quadrilateral, and so sin⁡C=sin⁡D\sin C = \sin D by the rule for the second quadrant. In both diagrams, therefore, sin⁡C=sin⁡D\sin C = \sin D, and hence

c=2Rsin⁡C,that iscsin⁡C=2R.c = 2R\sin C, \qquad\text{that is}\qquad \frac{c}{\sin C} = 2R .

Similarly bsin⁡B=2R\dfrac{b}{\sin B} = 2R and asin⁡A=2R\dfrac{a}{\sin A} = 2R, so

asin⁡A=bsin⁡B=csin⁡C=2R.\frac{a}{\sin A} = \frac{b}{\sin B} = \frac{c}{\sin C} = 2R .
OABCD(i) C and D on the same side of ABOABCD(ii) C and D on opposite sides
Figure 5.11. The circle through AA, BB and CC, with the diameter ADAD.

Any pair of these three equal ratios gives an equation containing two sides and two angles. So the sine rule can be used to solve a triangle in which we know either two sides and one angle, or two angles and one side, provided that one given side is opposite a given angle.

Example 5.11.

In a triangle ABCABC, A=73∘A = 73^\circ, B=49∘B = 49^\circ and a=12.2a = 12.2 cm. Find bb and cc.

As aa, AA and BB are known, bb can be calculated using

bsin⁡B=asin⁡A,b=12.2sin⁡49∘sin⁡73∘=9.63 cm.\frac{b}{\sin B} = \frac{a}{\sin A}, \qquad b = \frac{12.2\sin 49^\circ}{\sin 73^\circ} = 9.63 \text{ cm}.

Also C=180∘−73∘−49∘=58∘C = 180^\circ - 73^\circ - 49^\circ = 58^\circ, so cc can be calculated using

csin⁡C=asin⁡A,c=12.2sin⁡58∘sin⁡73∘=10.82 cm.\frac{c}{\sin C} = \frac{a}{\sin A}, \qquad c = \frac{12.2\sin 58^\circ}{\sin 73^\circ} = 10.82 \text{ cm}.

In this example the given data define one and only one triangle.

The Ambiguous Case

Sometimes, when two sides and one angle are specified, two different triangles can be found from the data. Suppose that we have to solve the triangle in which A=24∘A = 24^\circ, c=2.6c = 2.6 cm and a=1.1a = 1.1 cm. Knowing aa, AA and cc we can find CC using

asin⁡A=csin⁡C,sin⁡C=2.6sin⁡24∘1.1=0.9614.\frac{a}{\sin A} = \frac{c}{\sin C}, \qquad \sin C = \frac{2.6\sin 24^\circ}{1.1} = 0.9614 .

In a triangle, sin⁡C=0.9614\sin C = 0.9614 gives C=74∘C = 74^\circ or C=106∘C = 106^\circ, since the sine is positive in both the first and the second quadrants. We must check whether both are possible values for CC.

  1. If C=74∘C = 74^\circ, then A+C=98∘A + C = 98^\circ, which is less than 180∘180^\circ. So 74∘74^\circ is a possible value for CC, corresponding to B=82∘B = 82^\circ.
  2. If C=106∘C = 106^\circ, then A+C=130∘A + C = 130^\circ, which is also less than 180∘180^\circ. So 106∘106^\circ is a possible value too, corresponding to B=50∘B = 50^\circ.

So there are two possible triangles with the given data. This is known as the ambiguous case, and it is easily understood by attempting to construct the triangle from its specification: an arc of radius 1.11.1 about BB cuts the line from AA in two places. Each position of CC corresponds to a different pair of values for BB and bb, and in each case the solution of the triangle is completed using asin⁡A=bsin⁡B\dfrac{a}{\sin A} = \dfrac{b}{\sin B}, which gives b=2.68b = 2.68 cm and b=2.07b = 2.07 cm.

ABC1C224°2.61.11.1
Figure 5.12. The two positions C1C_1 and C2C_2 of the vertex CC in the ambiguous case.

There are not always two possible triangles when one angle and two sides are given, as the next example shows.

Example 5.12.

If A=37∘A = 37^\circ, a=4.59a = 4.59 cm and c=2.1c = 2.1 cm, show that there is only one possible triangle ABCABC, and find its remaining angles.

Using asin⁡A=csin⁡C\dfrac{a}{\sin A} = \dfrac{c}{\sin C},

sin⁡C=2.1sin⁡37∘4.59=0.2753,soC=16∘ or C=164∘.\sin C = \frac{2.1\sin 37^\circ}{4.59} = 0.2753, \qquad\text{so}\qquad C = 16^\circ \ \text{or}\ C = 164^\circ .

If C=16∘C = 16^\circ and A=37∘A = 37^\circ, then A+C=53∘A + C = 53^\circ, which is less than 180∘180^\circ, so 16∘16^\circ is a possible value for CC, corresponding to B=127∘B = 127^\circ. If C=164∘C = 164^\circ, then A+C=201∘A + C = 201^\circ, which is more than 180∘180^\circ, so 164∘164^\circ is not a possible value. Hence there is only one triangle defined by the given data, and its other angles are C=16∘C = 16^\circ and B=127∘B = 127^\circ.

Problem 5.4.

In each part the data refer to a triangle in the standard notation.

  1. B=35∘B = 35^\circ, a=2.7a = 2.7 cm and b=5.1b = 5.1 cm; find AA.
  2. b=3.8b = 3.8 cm, A=25∘A = 25^\circ and a=1.8a = 1.8 cm; find the possible values of CC.
  3. In a triangle PQRPQR the angle PQRPQR is 30∘30^\circ and the angle QPRQPR is θ\theta. Show that sin⁡θ=p2q\sin\theta = \dfrac{p}{2q}.

The Cosine Rule

The sine rule can be applied only when we know an angle opposite a given side. It cannot be used, for instance, when two sides and the included angle are given. Such cases are solved using the cosine rule,

a2=b2+c2−2bccos⁡A.a^2 = b^2 + c^2 - 2bc\cos A .

We show this by placing the triangle ABCABC on Cartesian axes with AA at the origin and ABAB along the positive xx-axis, as in Figure 5.13. In diagram (i) the coordinates of BB and CC are (c,0)(c, 0) and (bcos⁡A,bsin⁡A)(b\cos A, b\sin A). In diagram (ii), where AA is obtuse, they are (c,0)(c, 0) and (−bcos⁡(180∘−A),bsin⁡(180∘−A))\bigl(-b\cos(180^\circ - A), b\sin(180^\circ - A)\bigr), which by the rule for the second quadrant is again (bcos⁡A,bsin⁡A)(b\cos A, b\sin A). So by the length of the line joining two points,

a2=(bsin⁡A)2+(bcos⁡A−c)2=b2sin⁡2A+b2cos⁡2A−2bccos⁡A+c2,a^2 = (b\sin A)^2 + (b\cos A - c)^2 = b^2\sin^2 A + b^2\cos^2 A - 2bc\cos A + c^2,

where sin⁡2A\sin^2 A means (sin⁡A)2(\sin A)^2. Since CC is at a distance bb from the origin, (bcos⁡A)2+(bsin⁡A)2=b2(b\cos A)^2 + (b\sin A)^2 = b^2, and therefore

a2=b2+c2−2bccos⁡A.a^2 = b^2 + c^2 - 2bc\cos A .

Similarly it can be shown that

b2=c2+a2−2cacos⁡B,c2=a2+b2−2abcos⁡C.b^2 = c^2 + a^2 - 2ca\cos B, \qquad c^2 = a^2 + b^2 - 2ab\cos C .
xyAB (c, 0)Cabc(i) A acutexyAB (c, 0)Cabc(ii) A obtuse
Figure 5.13. The triangle ABCABC placed on axes, with AA acute and with AA obtuse.

The calculation involved in using the cosine rule is less straightforward than that required for the sine rule, so the cosine rule is used only when the sine rule is inapplicable.

Example 5.13.

In a triangle ABCABC, a=17.5a = 17.5 cm, b=8.4b = 8.4 cm and c=11.9c = 11.9 cm. Find the largest angle.

The longest side is opposite the largest angle, so we must find AA. As aa, bb and cc are given we use the cosine rule, rearranged so that AA can be found conveniently:

cos⁡A=b2+c2−a22bc=8.42+11.92−17.522(8.4)(11.9)=−0.4706.\cos A = \frac{b^2 + c^2 - a^2}{2bc} = \frac{8.4^2 + 11.9^2 - 17.5^2}{2(8.4)(11.9)} = -0.4706 .

Hence A=118.1∘A = 118.1^\circ, which is obtuse because cos⁡A\cos A is negative.

When the solution of a triangle requires the cosine rule first, the remaining sides and angles can then be found using the sine rule.

Example 5.14.

Solve the triangle PQRPQR given that p=12.1p = 12.1 cm, q=7.3q = 7.3 cm and R=37.5∘R = 37.5^\circ.

We first use

r2=p2+q2−2pqcos⁡R,which givesr=7.72 cm.r^2 = p^2 + q^2 - 2pq\cos R, \qquad\text{which gives}\qquad r = 7.72 \text{ cm}.

Now the sine rule can be used, since we know rr, RR and qq:

qsin⁡Q=rsin⁡R,sin⁡Q=7.3sin⁡37.5∘7.72=0.5756.\frac{q}{\sin Q} = \frac{r}{\sin R}, \qquad \sin Q = \frac{7.3\sin 37.5^\circ}{7.72} = 0.5756 .

Hence Q=35.1∘Q = 35.1^\circ. QQ cannot be obtuse, as qq is not the longest side of the triangle. Finally P=180∘−Q−R=107.4∘P = 180^\circ - Q - R = 107.4^\circ.

Problem 5.5.

  1. Solve the triangle in which a=4a = 4, b=7b = 7 and c=5c = 5.
  2. Using cos⁡A=b2+c2−a22bc\cos A = \dfrac{b^2 + c^2 - a^2}{2bc}, show that AA is acute if a2<b2+c2a^2 < b^2 + c^2 and obtuse if a2>b2+c2a^2 > b^2 + c^2, and check that taking A=90∘A = 90^\circ gives the result of Pythagoras.
  3. Find the angles of a triangle whose sides are in the ratio 2:3:42 : 3 : 4.

The Graphs of the Circular Functions

There is a single value for each trigonometric ratio of any angle, so the mappings θ↦sin⁡θ\theta \mapsto \sin\theta, θ↦cos⁡θ\theta \mapsto \cos\theta, and so on, are functions, and we can plot graphs showing how each trigonometric function behaves as θ\theta varies. The graphs of the sine, cosine and tangent are particularly important.

The Sine Function

−11θ−2π−3π/2−π−π/2π/2π3π/22π
Figure 5.14. The graph of f(θ)=sin⁡θf(\theta) = \sin\theta.

The graph of f:θ↦sin⁡θf : \theta \mapsto \sin\theta, θ∈R\theta \in \RR, shows that the sine function has the following characteristics.

  1. It is continuous: its graph has no breaks.
  2. Its range is −1⩽sin⁡θ⩽1-1 \leqslant \sin\theta \leqslant 1.
  3. The shape of the graph from θ=0\theta = 0 to θ=2π\theta = 2\pi is repeated for each further complete revolution.

Definition 5.15 (Periodic Function).

A function whose graph repeats a pattern is periodic, or cyclic. The width of the repeating pattern, measured on the horizontal axis, is the period of the function: it is the smallest positive number pp for which f(θ+p)=f(θ)f(\theta + p) = f(\theta) for every θ\theta.

So θ↦sin⁡θ\theta \mapsto \sin\theta is a periodic function with a period of 2π2\pi, a maximum value of 11 and a minimum value of −1-1. A graph of this shape is known as a sine wave. The amplitude is half the distance between the maximum and minimum values; for sin⁡θ\sin\theta its value is 11.

The Cosine Function

−11θ−2π−3π/2−π−π/2π/2π3π/22π
Figure 5.15. The graph of f(θ)=cos⁡θf(\theta) = \cos\theta, with the sine curve dashed.

The characteristics of the graph of f:θ↦cos⁡θf : \theta \mapsto \cos\theta, θ∈R\theta \in \RR, are as follows.

  1. It is continuous.
  2. It lies entirely within the range −1⩽cos⁡θ⩽1-1 \leqslant \cos\theta \leqslant 1.
  3. It is periodic with a period of 2π2\pi.
  4. It has the same shape as the sine graph, but is displaced a distance π2\dfrac{\pi}{2} to the left on the horizontal axis. Such a displacement is known as a phase difference, or phase shift.

So θ↦cos⁡θ\theta \mapsto \cos\theta is a cyclic function with period 2π2\pi and values from −1-1 to 11.

The Tangent Function

−3−2−1123θ−2π−3π/2−π−π/2π/2π3π/22π
Figure 5.16. The graph of f(θ)=tan⁡θf(\theta) = \tan\theta.

The behaviour of the tangent function f:θ↦tan⁡θf : \theta \mapsto \tan\theta is different from that of the sine and cosine functions in several respects.

  1. It is not continuous, being undefined when θ=±π2,±3π2,±5π2,…\theta = \pm\dfrac{\pi}{2}, \pm\dfrac{3\pi}{2}, \pm\dfrac{5\pi}{2}, \ldots, where the lines drawn dashed in Figure 5.16 are asymptotes to the curve.
  2. The range of possible values of tan⁡θ\tan\theta is unlimited.
  3. The tangent function is periodic, but its period is π\pi, not 2π2\pi as for the sine and cosine.

Special Values

It is useful to note the angles whose trigonometric ratios have the values 00 and ±1\pm 1, and the angles at which tan⁡θ\tan\theta is undefined. Reference to the graphs shows that, for n∈Zn \in \ZZ,

sin⁡θ=0when θ=…,−2π,−π,0,π,2π,3π,…,that is θ=nπ,sin⁡θ=1when θ=…,−3π2,π2,5π2,…,that is θ=2nπ+π2,sin⁡θ=−1when θ=…,−π2,3π2,7π2,…,that is θ=2nπ−π2,cos⁡θ=0when θ=…,−π2,π2,3π2,5π2,…,that is θ=(2n+1)π2,cos⁡θ=1when θ=…,−2π,0,2π,4π,…,that is θ=2nπ,cos⁡θ=−1when θ=…,−π,π,3π,5π,…,that is θ=(2n+1)π,tan⁡θ=0when θ=…,−π,0,π,2π,…,that is θ=nπ,\begin{aligned} \sin\theta &= 0 && \text{when } \theta = \ldots, -2\pi, -\pi, 0, \pi, 2\pi, 3\pi, \ldots, &&\text{that is } \theta = n\pi, \\ \sin\theta &= 1 && \text{when } \theta = \ldots, -\tfrac{3\pi}{2}, \tfrac{\pi}{2}, \tfrac{5\pi}{2}, \ldots, &&\text{that is } \theta = 2n\pi + \tfrac{\pi}{2}, \\ \sin\theta &= -1 && \text{when } \theta = \ldots, -\tfrac{\pi}{2}, \tfrac{3\pi}{2}, \tfrac{7\pi}{2}, \ldots, &&\text{that is } \theta = 2n\pi - \tfrac{\pi}{2}, \\ \cos\theta &= 0 && \text{when } \theta = \ldots, -\tfrac{\pi}{2}, \tfrac{\pi}{2}, \tfrac{3\pi}{2}, \tfrac{5\pi}{2}, \ldots, &&\text{that is } \theta = (2n + 1)\tfrac{\pi}{2}, \\ \cos\theta &= 1 && \text{when } \theta = \ldots, -2\pi, 0, 2\pi, 4\pi, \ldots, &&\text{that is } \theta = 2n\pi, \\ \cos\theta &= -1 && \text{when } \theta = \ldots, -\pi, \pi, 3\pi, 5\pi, \ldots, &&\text{that is } \theta = (2n + 1)\pi, \\ \tan\theta &= 0 && \text{when } \theta = \ldots, -\pi, 0, \pi, 2\pi, \ldots, &&\text{that is } \theta = n\pi, \end{aligned}

and tan⁡θ\tan\theta is undefined, its graph going off to ±∞\pm\infty, when θ=(2n+1)π2\theta = (2n + 1)\dfrac{\pi}{2}.

Remark.

Throughout, 2n2n stands for any even integer and 2n+12n + 1 for any odd integer, provided that n∈Zn \in \ZZ.

The Reciprocal Ratios

The three ratios sin⁡θ\sin\theta, cos⁡θ\cos\theta and tan⁡θ\tan\theta are used much more frequently than their reciprocals cosec⁡θ\operatorname{cosec}\theta, sec⁡θ\sec\theta and cot⁡θ\cot\theta, but the reciprocal ratios must not be overlooked. The graph of f(θ)=cosec⁡θf(\theta) = \operatorname{cosec}\theta can be drawn without any table of values, simply by observing the graph of f(θ)=sin⁡θf(\theta) = \sin\theta and using the following properties of any expression and its reciprocal.

  1. As an expression approaches zero its reciprocal becomes numerically large without limit, and as an expression becomes numerically large without limit its reciprocal approaches zero.
  2. The reciprocal of 11 is 11, and the reciprocal of −1-1 is −1-1.
  3. Where an expression has a maximum value its reciprocal has a minimum value, and conversely.
  4. Where an expression is increasing its reciprocal is decreasing, and conversely.
  5. An expression and its reciprocal have the same sign.

These properties are reasonable enough to accept without detailed analysis at this stage.

−3−2−1123θ−π−π/2π/2π3π/22π
Figure 5.17. The graph of f(θ)=cosec⁡θf(\theta) = \operatorname{cosec}\theta for −π⩽θ⩽2π-\pi \leqslant \theta \leqslant 2\pi, with the sine curve dashed.

In the same way the graphs of f(θ)=sec⁡θf(\theta) = \sec\theta and f(θ)=cot⁡θf(\theta) = \cot\theta can be deduced from those of cos⁡θ\cos\theta and tan⁡θ\tan\theta.

−3−2−1123θ−π−π/2π/2π3π/22πsec θ, with cos θ dashed−3−2−1123θ−π−π/2π/2π3π/22πcot θ, with tan θ dashed
Figure 5.18. The graphs of sec⁡θ\sec\theta and cot⁡θ\cot\theta, each drawn over the curve it is the reciprocal of.

Inverse Circular Functions

The function f:x↦sin⁡xf : x \mapsto \sin x is a many-one mapping for the domain x∈Rx \in \RR, and so it does not have an inverse function. However, if the domain is redefined as −π2⩽x⩽π2-\dfrac{\pi}{2} \leqslant x \leqslant \dfrac{\pi}{2}, the function f:x↦sin⁡xf : x \mapsto \sin x is a one-one mapping, and now it does have an inverse. This inverse sine function is denoted by arcsin⁡\arcsin or sin⁡−1\sin^{-1}. Thus

if f:x↦sin⁡x, −π2⩽x⩽π2,thenf−1:x↦arcsin⁡x, −1⩽x⩽1.\text{if } f : x \mapsto \sin x, \ -\frac{\pi}{2} \leqslant x \leqslant \frac{\pi}{2}, \qquad\text{then}\qquad f^{-1} : x \mapsto \arcsin x, \ -1 \leqslant x \leqslant 1 .

For f:x↦sin⁡xf : x \mapsto \sin x the input is an angle and the output is a number. So for the inverse function f−1:x↦arcsin⁡xf^{-1} : x \mapsto \arcsin x the input is a number and the output is an angle, and arcsin⁡x\arcsin x means “the angle whose sine is xx”.

Similarly, if f:x↦cos⁡xf : x \mapsto \cos x, 0⩽x⩽π0 \leqslant x \leqslant \pi, then f−1f^{-1} exists and is denoted by arccos⁡\arccos or cos⁡−1\cos^{-1}, where arccos⁡x\arccos x means “the angle whose cosine is xx”:

if f:x↦cos⁡x, 0⩽x⩽π,thenf−1:x↦arccos⁡x, −1⩽x⩽1.\text{if } f : x \mapsto \cos x, \ 0 \leqslant x \leqslant \pi, \qquad\text{then}\qquad f^{-1} : x \mapsto \arccos x, \ -1 \leqslant x \leqslant 1 .

Further, if f:x↦tan⁡xf : x \mapsto \tan x for −π2<x<π2-\dfrac{\pi}{2} < x < \dfrac{\pi}{2}, the inverse function f−1f^{-1} exists and is written arctan⁡\arctan or tan⁡−1\tan^{-1}, where arctan⁡x\arctan x means “the angle whose tangent is xx”. The domain of x↦arctan⁡xx \mapsto \arctan x is x∈Rx \in \RR. In the same way arccot⁡x\operatorname{arccot} x is the angle whose cotangent is xx, and for positive xx it is arctan⁡1x\arctan\dfrac{1}{x}.

−11−11xysin and arcsin−11π−11π/2πxycos and arccos−22−22xytan and arctan
Figure 5.19. The restricted sine, cosine and tangent functions and their inverses, reflected in the line y=xy = x.

Remark (A warning about notation).

If the notation sin⁡−1\sin^{-1} is adopted, it is most important to appreciate that sin⁡−1x\sin^{-1} x is not the same as 1sin⁡x\dfrac{1}{\sin x}.

Common Trigonometric Ratios

The angles 30∘30^\circ, 45∘45^\circ and 60∘60^\circ, and the angles for which these are the associated acute angles, such as 150∘150^\circ, 225∘225^\circ and 300∘300^\circ, are used frequently, so their trigonometric ratios are well worth noting.

Consider first an equilateral triangle ABCABC which is bisected by the line ADAD. In the triangle BADBAD the angle BB is 60∘60^\circ, or π3\dfrac{\pi}{3}, since the triangle ABCABC is equilateral, and the angle at AA is 30∘30^\circ, or π6\dfrac{\pi}{6}, since the angle BACBAC is bisected. If AB=2AB = 2 units then BD=1BD = 1 unit, and AD=3AD = \sqrt3 units by Pythagoras. Therefore

sin⁡π6=sin⁡30∘=12,sin⁡π3=sin⁡60∘=32,cos⁡π6=cos⁡30∘=32,cos⁡π3=cos⁡60∘=12,tan⁡π6=tan⁡30∘=13,tan⁡π3=tan⁡60∘=3.\begin{aligned} \sin\frac{\pi}{6} &= \sin 30^\circ = \frac12, & \sin\frac{\pi}{3} &= \sin 60^\circ = \frac{\sqrt3}{2}, \\ \cos\frac{\pi}{6} &= \cos 30^\circ = \frac{\sqrt3}{2}, & \cos\frac{\pi}{3} &= \cos 60^\circ = \frac12, \\ \tan\frac{\pi}{6} &= \tan 30^\circ = \frac{1}{\sqrt3}, & \tan\frac{\pi}{3} &= \tan 60^\circ = \sqrt3 . \end{aligned}

Alternatively we can say

arcsin⁡12=π6,arccos⁡32=π6,arctan⁡13=π6,arcsin⁡32=π3,arccos⁡12=π3,arctan⁡3=π3.\begin{aligned} \arcsin\frac12 &= \frac{\pi}{6}, & \arccos\frac{\sqrt3}{2} &= \frac{\pi}{6}, & \arctan\frac{1}{\sqrt3} &= \frac{\pi}{6}, \\ \arcsin\frac{\sqrt3}{2} &= \frac{\pi}{3}, & \arccos\frac12 &= \frac{\pi}{3}, & \arctan\sqrt3 &= \frac{\pi}{3} . \end{aligned}

Now consider a triangle ABCABC in which AB=BCAB = BC and the angle BB is a right angle, so that the angles at AA and CC are each 45∘45^\circ. If AB=BC=1AB = BC = 1 unit, then AC=2AC = \sqrt2 units by Pythagoras. Therefore

sin⁡π4=sin⁡45∘=12=22,cos⁡π4=cos⁡45∘=12=22,tan⁡π4=tan⁡45∘=1,\sin\frac{\pi}{4} = \sin 45^\circ = \frac{1}{\sqrt2} = \frac{\sqrt2}{2}, \qquad \cos\frac{\pi}{4} = \cos 45^\circ = \frac{1}{\sqrt2} = \frac{\sqrt2}{2}, \qquad \tan\frac{\pi}{4} = \tan 45^\circ = 1,

and arcsin⁡12=π4=arccos⁡12\arcsin\dfrac{1}{\sqrt2} = \dfrac{\pi}{4} = \arccos\dfrac{1}{\sqrt2}, arctan⁡1=π4\arctan 1 = \dfrac{\pi}{4}.

ABCD21√360°30°half of an equilateral triangleABC11√245°45°a right-angled isosceles triangle
Figure 5.20. The triangles which give the ratios of 30∘30^\circ, 60∘60^\circ and 45∘45^\circ.

Complementary Angles

Definition 5.16 (Complementary Angles).

If the sum of two acute angles is 90∘90^\circ, or π2\dfrac{\pi}{2}, they are complementary, and each is the complement of the other.

Consider a right-angled triangle ABCABC containing the angles α\alpha and β\beta, as in Figure 5.21. Then

sin⁡α=ac=cos⁡β,cos⁡α=bc=sin⁡β,tan⁡α=ab=cot⁡β,cot⁡α=ba=tan⁡β.\sin\alpha = \frac{a}{c} = \cos\beta, \qquad \cos\alpha = \frac{b}{c} = \sin\beta, \qquad \tan\alpha = \frac{a}{b} = \cot\beta, \qquad \cot\alpha = \frac{b}{a} = \tan\beta .
ACBαβabc
Figure 5.21. A right-angled triangle containing two complementary angles α\alpha and β\beta.

But α\alpha and β\beta are complementary, so we have shown that the sine of an angle is the cosine of its complement, and the tangent of an angle is the cotangent of its complement. Because of this property the sine and cosine of an angle are called complementary ratios, and similarly the tangent and cotangent are complementary ratios.

The same relationships hold for angles of any size. Reflecting OPOP in the line y=xy = x exchanges the coordinates xx and yy of PP, and turns the angle θ\theta into π2−θ\dfrac{\pi}{2} - \theta, so for every angle

sin⁡θ=cos⁡(π2−θ),cos⁡θ=sin⁡(π2−θ),tan⁡θ=cot⁡(π2−θ).\sin\theta = \cos\left(\frac{\pi}{2} - \theta\right), \qquad \cos\theta = \sin\left(\frac{\pi}{2} - \theta\right), \qquad \tan\theta = \cot\left(\frac{\pi}{2} - \theta\right) .

Trigonometric Equations

Definition 5.17 (Trigonometric Equation).

An equation in which at least one term contains a trigonometric ratio is a trigonometric equation. Solving it means finding the angle or angles for which it is true.

Consider the simple equation sin⁡θ=0\sin\theta = 0. Referring to the graph of the sine function, we see that sin⁡θ=0\sin\theta = 0 when θ\theta is any multiple of π\pi, that is when θ=nπ\theta = n\pi, where n∈Zn \in \ZZ. The full, or general, solution of this equation is the infinite set of angles θ=nπ\theta = n\pi, or θ=180n∘\theta = 180n^\circ.

Sometimes it is necessary to extract certain values of θ\theta from the infinite set. Solving the equation sin⁡θ=0\sin\theta = 0 for −π⩽θ⩽π-\pi \leqslant \theta \leqslant \pi, for instance, gives the finite solution set θ=−π,0,π\theta = -\pi, 0, \pi.

There are two basic approaches to finding the solution of a trigonometric equation. One of them was used above, and refers to the graph of the appropriate circular function; it is usually best for sines and cosines with the values ±1\pm 1 and 00, and for tangents which are zero or undefined. Alternatively, the position of the rotating line OPOP in the appropriate quadrants can lead to a clear solution. In all cases the first step is to find the principal solution, which is the principal value, PV, of θ\theta.

Principal Values

−11θ−π−π/2π/2πsin θ: principal values in [−π/2, π/2]−11θ−π−π/2π/2πcos θ: principal values in [0, π]−11θ−π−π/2π/2πtan θ: principal values between −π/2 and π/2
Figure 5.22. The intervals in which each value of the sine, cosine and tangent occurs exactly once.
  1. From the graph of the sine function, every possible value of sin⁡θ\sin\theta occurs once and only once in the interval −π2⩽θ⩽π2-\dfrac{\pi}{2} \leqslant \theta \leqslant \dfrac{\pi}{2}. So any equation sin⁡θ=s\sin\theta = s, with −1⩽s⩽1-1 \leqslant s \leqslant 1, has one and only one solution in this interval, and this is the principal value of θ\theta, in either the first or the fourth quadrant. For example, if sin⁡θ=12\sin\theta = \frac12 the principal solution is θ=π6\theta = \dfrac{\pi}{6}, and if sin⁡θ=−12\sin\theta = -\frac12 it is θ=−π6\theta = -\dfrac{\pi}{6}.
  2. Every possible value of cos⁡θ\cos\theta occurs once and only once in the interval 0⩽θ⩽π0 \leqslant \theta \leqslant \pi, so there is one and only one solution of cos⁡θ=c\cos\theta = c in this interval. This is the principal value of θ\theta, in either the first or the second quadrant. If cos⁡θ=12\cos\theta = \frac12 the principal solution is θ=π3\theta = \dfrac{\pi}{3}, and if cos⁡θ=−12\cos\theta = -\frac12 it is θ=2π3\theta = \dfrac{2\pi}{3}.
  3. Every possible value of tan⁡θ\tan\theta occurs once and only once for angles in the interval −π2<θ<π2-\dfrac{\pi}{2} < \theta < \dfrac{\pi}{2}, so one and only one solution of tan⁡θ=t\tan\theta = t is in this interval, and it is the principal value of θ\theta, in the first or the fourth quadrant. If tan⁡θ=1\tan\theta = 1 the principal solution is θ=π4\theta = \dfrac{\pi}{4}, and if tan⁡θ=−1\tan\theta = -1 it is θ=−π4\theta = -\dfrac{\pi}{4}.

These intervals are the ranges of the inverse functions, so the principal values of the solutions of sin⁡θ=s\sin\theta = s, cos⁡θ=c\cos\theta = c and tan⁡θ=t\tan\theta = t are arcsin⁡s\arcsin s, arccos⁡c\arccos c and arctan⁡t\arctan t.

Secondary Values

Having found the principal value of the solution of a trigonometric equation, we usually find a second angle with the same trigonometric ratio in the interval −π<θ⩽π-\pi < \theta \leqslant \pi. This solution lies in a different quadrant and is called the secondary value, SV, of θ\theta, or the secondary solution of the equation.

  1. If sin⁡θ=12\sin\theta = \frac12, the secondary solution is in the second quadrant, where the sine ratio is also positive, and it is θ=5π6\theta = \dfrac{5\pi}{6}. If sin⁡θ=−12\sin\theta = -\frac12, the secondary solution is in the third quadrant, where the sine ratio is also negative, and it is θ=−5π6\theta = -\dfrac{5\pi}{6}.
  2. If cos⁡θ=12\cos\theta = \frac12, the secondary value is in the fourth quadrant and is θ=−π3\theta = -\dfrac{\pi}{3}. If cos⁡θ=−12\cos\theta = -\frac12, the secondary value is in the third quadrant and is θ=−2π3\theta = -\dfrac{2\pi}{3}. For an equation of the form cos⁡θ=c\cos\theta = c, SV=−PV\text{SV} = -\text{PV}.
  3. If tan⁡θ=1\tan\theta = 1, the secondary solution is in the third quadrant and is θ=−3π4\theta = -\dfrac{3\pi}{4}. If tan⁡θ=−1\tan\theta = -1, the secondary value is in the second quadrant and is θ=3π4\theta = \dfrac{3\pi}{4}.
PV π/6SV 5π/6sin θ = ½PV π/3SV −π/3cos θ = ½PV π/4SV −3π/4tan θ = 1PV −π/6SV −5π/6sin θ = −½PV 2π/3SV −2π/3cos θ = −½PV −π/4SV 3π/4tan θ = −1
Figure 5.23. Principal and secondary values for six equations.

Problem 5.6.

  1. Determine the principal solutions of sin⁡θ=−32\sin\theta = -\dfrac{\sqrt3}{2}, cos⁡θ=−12\cos\theta = -\dfrac{1}{\sqrt2} and tan⁡θ=−33\tan\theta = -\dfrac{\sqrt3}{3}.
  2. Find the principal and secondary solutions of sin⁡θ=12\sin\theta = \dfrac{1}{\sqrt2}, cos⁡θ=−32\cos\theta = -\dfrac{\sqrt3}{2} and tan⁡θ=−3\tan\theta = -\sqrt3.
  3. By referring to the graphs of the appropriate circular functions, explain why the equations sin⁡θ=−1\sin\theta = -1 and cos⁡θ=1\cos\theta = 1 have no secondary solution.

Solutions in a Specified Range

In solving a trigonometric equation in a specified range we first find the principal angle and the secondary angle, except in those cases where there is no secondary angle. A quadrant diagram can then be drawn showing the two solution positions, and any angle measured from the positive xx-axis to either of the solution positions is a solution of the equation.

Example 5.18.

Solve the equation sin⁡θ=0.4\sin\theta = 0.4 within the interval −360∘⩽θ⩽360∘-360^\circ \leqslant \theta \leqslant 360^\circ.

The principal solution is θ=23.58∘\theta = 23.58^\circ. The secondary solution is in the second quadrant, since the sine ratio is positive in the first and second quadrants, and it is θ=156.42∘\theta = 156.42^\circ. Therefore the solutions within the specified interval are

θ=−336.42∘,−203.58∘,23.58∘,156.42∘.\theta = -336.42^\circ, \quad -203.58^\circ, \quad 23.58^\circ, \quad 156.42^\circ .
−360°−270°−180°−90°90°180°270°360°θ−110.4
Figure 5.24. The four solutions of sin⁡θ=0.4\sin\theta = 0.4 between −360∘-360^\circ and 360∘360^\circ.

Example 5.19.

Solve the equation tan⁡θ=−13\tan\theta = -\dfrac{1}{\sqrt3} in the interval 0⩽θ⩽2π0 \leqslant \theta \leqslant 2\pi.

If tan⁡θ=−13\tan\theta = -\dfrac{1}{\sqrt3}, the principal solution is in the fourth quadrant and the secondary solution is in the second quadrant: the PV is −π6-\dfrac{\pi}{6} and the SV is 5π6\dfrac{5\pi}{6}. Within the specified interval the solution set is

θ=5π6,11π6.\theta = \frac{5\pi}{6}, \quad \frac{11\pi}{6} .

As here, the principal value is not always included in the solution set.

Example 5.20.

Find the angles in the interval −360∘⩽θ⩽0-360^\circ \leqslant \theta \leqslant 0 which satisfy the equation cos⁡θ=0.7\cos\theta = 0.7.

Since cos⁡θ\cos\theta is positive, the principal solution is in the first quadrant and the secondary solution is in the fourth: the PV is 45.57∘45.57^\circ and the SV is −45.57∘-45.57^\circ. In the interval −360∘⩽θ⩽0-360^\circ \leqslant \theta \leqslant 0 the solution set is θ=−314.43∘,−45.57∘\theta = -314.43^\circ, -45.57^\circ.

Example 5.21.

Solve, within the interval 0⩽θ⩽360∘0 \leqslant \theta \leqslant 360^\circ, the equation sin⁡θ+3sin⁡θcos⁡θ=0\sin\theta + 3\sin\theta\cos\theta = 0.

First the equation must be factorised:

sin⁡θ (1+3cos⁡θ)=0.\sin\theta\,(1 + 3\cos\theta) = 0 .

Therefore either sin⁡θ=0\sin\theta = 0 or cos⁡θ=−13\cos\theta = -\frac13.

  1. Referring to the sine graph, sin⁡θ=0\sin\theta = 0 gives θ=0,180∘,360∘\theta = 0, 180^\circ, 360^\circ.
  2. For cos⁡θ=−13\cos\theta = -\frac13 the principal solution is in the second quadrant and the secondary solution in the third: the PV is 109.47∘109.47^\circ and the SV is −109.47∘-109.47^\circ. Within the specified range these give θ=109.47∘,250.53∘\theta = 109.47^\circ, 250.53^\circ.

So the complete solution set from 00 to 360∘360^\circ is θ=0,109.47∘,180∘,250.53∘,360∘\theta = 0, 109.47^\circ, 180^\circ, 250.53^\circ, 360^\circ.

Although the solutions of sin⁡θ=0\sin\theta = 0 could conveniently be expressed in radians, degrees are used because the range is specified in degrees. Units must not be mixed in any one example.

Problem 5.7.

Solve the following equations for angles in the interval 0⩽θ⩽2π0 \leqslant \theta \leqslant 2\pi.

  1. 3tan⁡θ=2sin⁡θ\sqrt3\tan\theta = 2\sin\theta;
  2. 2sin⁡θcos⁡θ+sin⁡θ=02\sin\theta\cos\theta + \sin\theta = 0;
  3. 4cos⁡θ=cos⁡θcosec⁡θ4\cos\theta = \cos\theta\operatorname{cosec}\theta.

General Solutions

Definition 5.22 (General Solution).

The general solution of a trigonometric equation is an expression which represents all the angles which satisfy the equation, that is an infinite set of angles.

In looking for a general solution we use the graphs of the circular functions, the period of each circular function, and the principal solution together with, except when the tangent ratio is involved, the secondary solution.

Consider the equation sin⁡θ=s\sin\theta = s, where −1⩽s⩽1-1 \leqslant s \leqslant 1. The period 2π2\pi of the sine function is covered by the interval −π<θ⩽π-\pi < \theta \leqslant \pi, which includes both the PV and the SV of θ\theta. So by adding or subtracting any multiple of 2π2\pi to either the PV or the SV we get another angle with the same sine. Thus the complete solution of sin⁡θ=s\sin\theta = s is

θ=PV+2nπorθ=SV+2nπ,n∈Z,\theta = \text{PV} + 2n\pi \quad\text{or}\quad \theta = \text{SV} + 2n\pi, \qquad n \in \ZZ,

or, in degrees, θ=PV+360n∘\theta = \text{PV} + 360n^\circ or θ=SV+360n∘\theta = \text{SV} + 360n^\circ.

A similar situation arises for the equation cos⁡θ=c\cos\theta = c, because both the PV and the SV of the cosine lie within one period, which is again 2π2\pi. So the complete solution of cos⁡θ=c\cos\theta = c is also given by adding multiples of 2π2\pi to either the PV or the SV. Remembering that for cosines the PV and the SV are equal in value but opposite in sign, the general solution of cos⁡θ=c\cos\theta = c can be given in the form

θ=±PV+2nπorθ=±PV+360n∘,n∈Z.\theta = \pm\text{PV} + 2n\pi \quad\text{or}\quad \theta = \pm\text{PV} + 360n^\circ, \qquad n \in \ZZ .

For the equation tan⁡θ=t\tan\theta = t, only the principal value is included in the complete period −π2<θ<π2-\dfrac{\pi}{2} < \theta < \dfrac{\pi}{2}. All further angles with the same tangent are given by adding multiples of π\pi, the period, to the PV:

θ=PV+nπ,n∈Z.\theta = \text{PV} + n\pi, \qquad n \in \ZZ .

Example 5.23.

Find the general solution of each of the equations tan⁡θ=1\tan\theta = 1 and tan⁡θ=−3\tan\theta = -\sqrt3.

The principal solution of tan⁡θ=1\tan\theta = 1 is θ=π4\theta = \dfrac{\pi}{4}, so the general solution is θ=π4+nπ\theta = \dfrac{\pi}{4} + n\pi, where n∈Zn \in \ZZ.

If tan⁡θ=−3\tan\theta = -\sqrt3, the principal solution is θ=−π3\theta = -\dfrac{\pi}{3}, and the general solution is therefore θ=nπ−π3\theta = n\pi - \dfrac{\pi}{3}.

Example 5.24.

Find the general solution of each of the equations cos⁡θ=12\cos\theta = \dfrac{1}{\sqrt2} and cos⁡θ=−12\cos\theta = -\dfrac12.

The principal solution of cos⁡θ=12\cos\theta = \dfrac{1}{\sqrt2} is θ=π4\theta = \dfrac{\pi}{4}, so the general solution is θ=2nπ±π4\theta = 2n\pi \pm \dfrac{\pi}{4}.

When cos⁡θ=−12\cos\theta = -\frac12, the principal value of θ\theta is 2π3\dfrac{2\pi}{3}, and the general solution is θ=2nπ±2π3\theta = 2n\pi \pm \dfrac{2\pi}{3}.

Example 5.25.

Find the general solution set of the equation sin⁡θ=12\sin\theta = \frac12.

The principal value of θ\theta for which sin⁡θ=12\sin\theta = \frac12 is π6\dfrac{\pi}{6}, and the secondary value, in the second quadrant, is 5π6\dfrac{5\pi}{6}. So the general solution set includes

θ=π6+2nπandθ=5π6+2nπ.\theta = \frac{\pi}{6} + 2n\pi \qquad\text{and}\qquad \theta = \frac{5\pi}{6} + 2n\pi .

Remark.

There is a way of combining the two parts of the general solution of sin⁡θ=s\sin\theta = s into one formula,

θ=(−1)n PV+nπ,n∈Z.\theta = (-1)^n\,\text{PV} + n\pi, \qquad n \in \ZZ .

When nn is even this is PV+2kπ\text{PV} + 2k\pi, and when nn is odd it is π−PV+2kπ\pi - \text{PV} + 2k\pi, which is the secondary value plus a multiple of 2π2\pi. So when sin⁡θ=12\sin\theta = \frac12, θ=(−1)nπ6+nπ\theta = (-1)^n\dfrac{\pi}{6} + n\pi. Either form of the general solution may be used.

Example 5.26.

Find the general solution of the equation 4sin⁡θ (2tan⁡θ+3)+6tan⁡θ+9=04\sin\theta\,(2\tan\theta + 3) + 6\tan\theta + 9 = 0.

First we simplify and factorise the equation:

4sin⁡θ (2tan⁡θ+3)+3(2tan⁡θ+3)=0,(4sin⁡θ+3)(2tan⁡θ+3)=0.4\sin\theta\,(2\tan\theta + 3) + 3(2\tan\theta + 3) = 0, \qquad (4\sin\theta + 3)(2\tan\theta + 3) = 0 .

So either sin⁡θ=−34\sin\theta = -\frac34 or tan⁡θ=−32\tan\theta = -\frac32.

  1. For sin⁡θ=−34\sin\theta = -\frac34 the principal solution is θ=−48.59∘\theta = -48.59^\circ, and the secondary solution, in the third quadrant, is θ=−131.41∘\theta = -131.41^\circ. Hence the general solution is θ=−48.59∘+360n∘\theta = -48.59^\circ + 360n^\circ or θ=−131.41∘+360n∘\theta = -131.41^\circ + 360n^\circ.
  2. For tan⁡θ=−32\tan\theta = -\frac32 the principal solution is the only one we need, and it is θ=−56.31∘\theta = -56.31^\circ. The general solution is then θ=−56.31∘+180n∘\theta = -56.31^\circ + 180n^\circ.

Combining these results, the general solution set of the given equation is

θ=−48.59∘+360n∘,θ=−131.41∘+360n∘,θ=−56.31∘+180n∘.\theta = -48.59^\circ + 360n^\circ, \qquad \theta = -131.41^\circ + 360n^\circ, \qquad \theta = -56.31^\circ + 180n^\circ .

Problem 5.8.

Find the general solution of each of the following equations.

  1. sin⁡θ=−32\sin\theta = -\dfrac{\sqrt3}{2};
  2. cos⁡θ=0\cos\theta = 0;
  3. cos⁡θ=0.371\cos\theta = 0.371.

Multiple Angles

Equations are frequently met in which the angle involved is a multiple of θ\theta, such as cos⁡2θ=12\cos 2\theta = \frac12 or tan⁡3θ=−2\tan 3\theta = -2. Such equations are solved by determining first the necessary values of the multiple angle and then, by division, the corresponding values of θ\theta.

Example 5.27.

Find the general solution of the equation cos⁡2θ=12\cos 2\theta = \frac12.

Let 2θ=φ2\theta = \varphi, so that cos⁡φ=12\cos\varphi = \frac12. The principal value of φ\varphi is π3\dfrac{\pi}{3}, so the general solution for φ\varphi is φ=±π3+2nπ\varphi = \pm\dfrac{\pi}{3} + 2n\pi, that is

2θ=±π3+2nπ,henceθ=±π6+nπ.2\theta = \pm\frac{\pi}{3} + 2n\pi, \qquad\text{hence}\qquad \theta = \pm\frac{\pi}{6} + n\pi .

Example 5.28.

Find the angles within the range −180∘⩽θ⩽180∘-180^\circ \leqslant \theta \leqslant 180^\circ which satisfy the equation tan⁡3θ=−2\tan 3\theta = -2.

Let 3θ=φ3\theta = \varphi, so that tan⁡φ=−2\tan\varphi = -2. The principal value is in the fourth quadrant, φ=−63.43∘\varphi = -63.43^\circ, and the secondary value is in the second quadrant, φ=116.57∘\varphi = 116.57^\circ. Values of θ\theta are required in the range −180∘⩽θ⩽180∘-180^\circ \leqslant \theta \leqslant 180^\circ, and φ=3θ\varphi = 3\theta, so we need the values of φ\varphi in the range −540∘⩽φ⩽540∘-540^\circ \leqslant \varphi \leqslant 540^\circ. These are

φ=−423.43∘,−243.43∘,−63.43∘,116.57∘,296.57∘,476.57∘,\varphi = -423.43^\circ, \quad -243.43^\circ, \quad -63.43^\circ, \quad 116.57^\circ, \quad 296.57^\circ, \quad 476.57^\circ,

and therefore

θ=−141.14∘,−81.14∘,−21.14∘,38.86∘,98.86∘,158.86∘.\theta = -141.14^\circ, \quad -81.14^\circ, \quad -21.14^\circ, \quad 38.86^\circ, \quad 98.86^\circ, \quad 158.86^\circ .

Alternatively, quoting the general solution for φ\varphi gives φ=−63.43∘+180n∘\varphi = -63.43^\circ + 180n^\circ, so that θ=−21.14∘+60n∘\theta = -21.14^\circ + 60n^\circ. Giving nn the values −2,−1,0,1,2,3-2, -1, 0, 1, 2, 3, which cover the required range, gives the same values of θ\theta.

Example 5.29.

Find the solutions of the equation sin⁡θ2=0.6\sin\dfrac{\theta}{2} = 0.6 for values of θ\theta between 00 and 360∘360^\circ.

Let θ2=φ\dfrac{\theta}{2} = \varphi, so that sin⁡φ=0.6\sin\varphi = 0.6. The principal value of φ\varphi is 36.87∘36.87^\circ, and the secondary value, in the second quadrant, is 143.13∘143.13^\circ. The required range of values of θ\theta is from 00 to 360∘360^\circ, so the range of values of φ\varphi is from 00 to 180∘180^\circ, which gives φ=36.87∘,143.13∘\varphi = 36.87^\circ, 143.13^\circ. Therefore θ=73.74∘,286.26∘\theta = 73.74^\circ, 286.26^\circ.

Problem 5.9.

  1. Solve the equations tan⁡2θ=1\tan 2\theta = 1 and sin⁡3θ=0.7\sin 3\theta = 0.7 within the interval 0⩽θ⩽360∘0 \leqslant \theta \leqslant 360^\circ.
  2. Find the general solution of the equation cos⁡2θ=0.63\cos 2\theta = 0.63.

The Equation cos⁡A=cos⁡B\cos A = \cos B

This type of equation can be solved very neatly as follows. Let cos⁡A=cos⁡B=c\cos A = \cos B = c, where −1⩽c⩽1-1 \leqslant c \leqslant 1. In general there are two solution positions for cos⁡B=c\cos B = c, OP1OP_1 and OP2OP_2, and the set of angles represented by OP1OP_1 and OP2OP_2 is 2nπ±B2n\pi \pm B. But we also know that cos⁡A=c\cos A = c, so OP1OP_1 and OP2OP_2 together represent all possible values of AA. Thus

cos⁡A=cos⁡BgivesA=2nπ±B,\cos A = \cos B \qquad\text{gives}\qquad A = 2n\pi \pm B,

that is, the values of AA are the general solution set for BB. The same argument for tangents and sines gives

tan⁡A=tan⁡B  gives  A=nπ+B,sin⁡A=sin⁡B  gives  A=2nπ+B  or  A=(2n+1)π−B.\tan A = \tan B \ \text{ gives } \ A = n\pi + B, \qquad \sin A = \sin B \ \text{ gives } \ A = 2n\pi + B \ \text{ or } \ A = (2n + 1)\pi - B .

Example 5.30.

Solve the equation cos⁡4θ=cos⁡θ\cos 4\theta = \cos\theta.

Using the conclusion above,

4θ=2nπ±θ,so5θ=2nπor3θ=2nπ,4\theta = 2n\pi \pm \theta, \qquad\text{so}\qquad 5\theta = 2n\pi \quad\text{or}\quad 3\theta = 2n\pi,

and hence θ=2nπ5\theta = \dfrac{2n\pi}{5} or θ=2nπ3\theta = \dfrac{2n\pi}{3}.

Example 5.31.

Find the values in the range 0⩽θ⩽360∘0 \leqslant \theta \leqslant 360^\circ which satisfy the equation tan⁡(3θ−40∘)=tan⁡θ\tan(3\theta - 40^\circ) = \tan\theta.

The general solution is

3θ−40∘=180n∘+θ,2θ=180n∘+40∘,θ=90n∘+20∘.3\theta - 40^\circ = 180n^\circ + \theta, \qquad 2\theta = 180n^\circ + 40^\circ, \qquad \theta = 90n^\circ + 20^\circ .

For 0⩽θ⩽360∘0 \leqslant \theta \leqslant 360^\circ we let n=0,1,2,3n = 0, 1, 2, 3, giving θ=20∘,110∘,200∘,290∘\theta = 20^\circ, 110^\circ, 200^\circ, 290^\circ.

This method can be used only for equations containing two terms involving the same trigonometric ratio. That situation can sometimes be arranged in an apparently unsuitable case, as in the next example.

Example 5.32.

Find the general solution of cos⁡3θ=sin⁡θ\cos 3\theta = \sin\theta.

We know that sin⁡θ=cos⁡(π2−θ)\sin\theta = \cos\left(\dfrac{\pi}{2} - \theta\right), from complementary angles, so the equation can be written

cos⁡3θ=cos⁡(π2−θ),so3θ=2nπ±(π2−θ).\cos 3\theta = \cos\left(\frac{\pi}{2} - \theta\right), \qquad\text{so}\qquad 3\theta = 2n\pi \pm \left(\frac{\pi}{2} - \theta\right) .

Therefore either 4θ=2nπ+π24\theta = 2n\pi + \dfrac{\pi}{2} or 2θ=2nπ−π22\theta = 2n\pi - \dfrac{\pi}{2}, that is

θ=nπ2+π8orθ=nπ−π4.\theta = \frac{n\pi}{2} + \frac{\pi}{8} \qquad\text{or}\qquad \theta = n\pi - \frac{\pi}{4} .

Problem 5.10.

  1. Find the general solutions of the equations cos⁡4θ=cos⁡3θ\cos 4\theta = \cos 3\theta and sin⁡4θ=sin⁡3θ\sin 4\theta = \sin 3\theta.
  2. Solve the equation cos⁡(2θ+60∘)=cos⁡θ\cos(2\theta + 60^\circ) = \cos\theta, giving the values of θ\theta from −180∘-180^\circ to 180∘180^\circ.

Graphs of Multiple and Compound Angles

Consider f:θ↦sin⁡2θf : \theta \mapsto \sin 2\theta. The following table gives pairs of corresponding values of θ\theta and f(θ)f(\theta).

θ\theta00π4\frac{\pi}{4}π2\frac{\pi}{2}3π4\frac{3\pi}{4}π\pi5π4\frac{5\pi}{4}3π2\frac{3\pi}{2}7π4\frac{7\pi}{4}2π2\pi
2θ2\theta00π2\frac{\pi}{2}π\pi3π2\frac{3\pi}{2}2π2\pi5π2\frac{5\pi}{2}3π3\pi7π2\frac{7\pi}{2}4π4\pi
f(θ)f(\theta)001100−1-1001100−1-100
−11θπ/4π/23π/4π5π/43π/27π/42π
Figure 5.25. The graph of f(θ)=sin⁡2θf(\theta) = \sin 2\theta, with the graph of sin⁡θ\sin\theta dashed.

The following characteristics can be observed.

  1. The function sin⁡2θ\sin 2\theta is cyclic, and its period is π\pi, that is 12×2π\frac12 \times 2\pi.
  2. Its range is −1⩽sin⁡2θ⩽1-1 \leqslant \sin 2\theta \leqslant 1.
  3. Its shape is a sine wave.
  4. Within the domain 0⩽θ⩽2π0 \leqslant \theta \leqslant 2\pi there are two complete cycles of the curve, compared with only one for the basic function θ↦sin⁡θ\theta \mapsto \sin\theta: the complete cycle appears with twice the frequency.

When the same investigation is carried out on θ↦sin⁡3θ\theta \mapsto \sin 3\theta, we find that the function is cyclic with period 2π3\dfrac{2\pi}{3}, so that three complete cycles occur between 00 and 2π2\pi. In general the graph of θ↦sin⁡kθ\theta \mapsto \sin k\theta, for k>0k > 0, is a sine wave with period 2πk\dfrac{2\pi}{k} and a frequency kk times that of θ↦sin⁡θ\theta \mapsto \sin\theta. We show this by noting that

sin⁡k(θ+2πk)=sin⁡(kθ+2π)=sin⁡kθ,\sin k\left(\theta + \frac{2\pi}{k}\right) = \sin(k\theta + 2\pi) = \sin k\theta,

so that the graph repeats after a width of 2πk\dfrac{2\pi}{k}, while as θ\theta runs through any interval of that width, kθk\theta runs through an interval of width 2π2\pi and so through one complete sine wave. Similar properties hold for θ↦cos⁡kθ\theta \mapsto \cos k\theta, whose period is 2πk\dfrac{2\pi}{k}, and for θ↦tan⁡kθ\theta \mapsto \tan k\theta, whose period is πk\dfrac{\pi}{k}. For example, the period of tan⁡2θ\tan 2\theta is π2\dfrac{\pi}{2} and its frequency is 22; the period of cos⁡4θ\cos 4\theta is π2\dfrac{\pi}{2} and its frequency is 44; the period of sin⁡θ2\sin\dfrac{\theta}{2} is 4π4\pi and its frequency is 12\frac12.

−22θπ/4π/23π/4π5π/43π/27π/42πtan 2θ, with tan θ dashed−11θπ/4π/23π/4π5π/43π/27π/42πcos 4θ, with cos θ dashed−11θπ/2π3π/22π5π/23π7π/24πsin ½θ, with sin θ dashed
Figure 5.26. The graphs of tan⁡2θ\tan 2\theta, cos⁡4θ\cos 4\theta and sin⁡12θ\sin\frac12\theta.

Now consider the function f(θ)=cos⁡(θ−α)f(\theta) = \cos(\theta - \alpha). This is clearly a cosine function, but

  1. f(θ)=0f(\theta) = 0 when θ−α=π2,3π2,…\theta - \alpha = \dfrac{\pi}{2}, \dfrac{3\pi}{2}, \ldots, that is when θ=π2+α,3π2+α,…\theta = \dfrac{\pi}{2} + \alpha, \dfrac{3\pi}{2} + \alpha, \ldots;
  2. f(θ)=1f(\theta) = 1 when θ−α=0,2π,4π,…\theta - \alpha = 0, 2\pi, 4\pi, \ldots, that is when θ=α,2π+α,4π+α,…\theta = \alpha, 2\pi + \alpha, 4\pi + \alpha, \ldots;
  3. f(θ)=−1f(\theta) = -1 when θ−α=π,3π,5π,…\theta - \alpha = \pi, 3\pi, 5\pi, \ldots, that is when θ=π+α,3π+α,5π+α,…\theta = \pi + \alpha, 3\pi + \alpha, 5\pi + \alpha, \ldots.

So the graph of f(θ)=cos⁡(θ−π4)f(\theta) = \cos\left(\theta - \dfrac{\pi}{4}\right), for example, is identical in shape to the graph of f(θ)=cos⁡θf(\theta) = \cos\theta, but is in a position given by moving the standard cosine curve a horizontal distance π4\dfrac{\pi}{4} to the right. Similarly the graph of f(θ)=cos⁡(θ+α)f(\theta) = \cos(\theta + \alpha) is given by moving the standard cosine curve a horizontal distance α\alpha to the left, as the graph of cos⁡(θ+π3)\cos\left(\theta + \dfrac{\pi}{3}\right) shows.

−11θ−π−3π/4−π/2−π/4π/4π/23π/4π5π/43π/27π/42πcos(θ − π/4), with cos θ dashed−11θ−π−2π/3−π/3π/32π/3π4π/35π/32πcos(θ + π/3), with cos θ dashed
Figure 5.27. The graphs of cos⁡(θ−π4)\cos\left(\theta - \frac{\pi}{4}\right) and cos⁡(θ+π3)\cos\left(\theta + \frac{\pi}{3}\right).

Now consider the function f(θ)=sin⁡(2θ+α)f(\theta) = \sin(2\theta + \alpha). Adding a constant angle moves a curve to the left, and in this case sin⁡(2θ+α)=0\sin(2\theta + \alpha) = 0 when θ=−α2\theta = -\dfrac{\alpha}{2}, so the graph of this function is obtained by moving the graph of sin⁡2θ\sin 2\theta a distance α2\dfrac{\alpha}{2} to the left. For example, the graph of sin⁡(2θ+π2)\sin\left(2\theta + \dfrac{\pi}{2}\right) is that of sin⁡2θ\sin 2\theta moved π4\dfrac{\pi}{4} to the left.

Similarly, if f(θ)=cos⁡(3θ−π4)f(\theta) = \cos\left(3\theta - \dfrac{\pi}{4}\right), putting 3θ−π4=03\theta - \dfrac{\pi}{4} = 0 shows that the graph is obtained by moving the graph of cos⁡3θ\cos 3\theta a distance π12\dfrac{\pi}{12} to the right. And if f(θ)=tan⁡(θ2+π4)f(\theta) = \tan\left(\dfrac{\theta}{2} + \dfrac{\pi}{4}\right), we have a tangent curve with a frequency of 12\frac12 which is moved a distance π2\dfrac{\pi}{2} to the left.

−11θ−π/2−π/4π/4π/23π/4π5π/43π/27π/42πsin(2θ + π/2), with sin 2θ dashed−11θπ/6π/3π/22π/35π/6π7π/64π/33π/25π/311π/62πcos(3θ − π/4), with cos 3θ dashed−22θ−2π−3π/2−π−π/2π/2π3π/22πtan(½θ + π/4), with tan ½θ dashed
Figure 5.28. The graphs of sin⁡(2θ+π2)\sin\left(2\theta + \frac{\pi}{2}\right), cos⁡(3θ−π4)\cos\left(3\theta - \frac{\pi}{4}\right) and tan⁡(12θ+π4)\tan\left(\frac12\theta + \frac{\pi}{4}\right).

An equation containing a compound angle is solved in the same way as one containing a multiple angle.

Example 5.33.

Find the general solution of the equation cos⁡(2θ−π6)=12\cos\left(2\theta - \dfrac{\pi}{6}\right) = \dfrac12.

The principal value of 2θ−π62\theta - \dfrac{\pi}{6} is π3\dfrac{\pi}{3}, so

2θ−π6=±π3+2nπ,2θ=±π3+2nπ+π6,θ=±π6+nπ+π12,2\theta - \frac{\pi}{6} = \pm\frac{\pi}{3} + 2n\pi, \qquad 2\theta = \pm\frac{\pi}{3} + 2n\pi + \frac{\pi}{6}, \qquad \theta = \pm\frac{\pi}{6} + n\pi + \frac{\pi}{12},

that is θ=nπ+π4\theta = n\pi + \dfrac{\pi}{4} or θ=nπ−π12\theta = n\pi - \dfrac{\pi}{12}.

Problem 5.11.

  1. Sketch the graphs of sin⁡4θ\sin 4\theta and sec⁡2θ\sec 2\theta in the domain 0⩽θ⩽2π0 \leqslant \theta \leqslant 2\pi, and state the period and the frequency of each function.
  2. Find the general solution of the equation cos⁡(θ−π6)=−12\cos\left(\theta - \dfrac{\pi}{6}\right) = -\dfrac12.

Trigonometric Identities

Any one angle has six trigonometric ratios, and a particular value of one ratio applies to an infinite set of angles, so it is not surprising that relationships exist between the various circular functions. They are identities, true for every angle for which the ratios are defined, and they are very useful in the development of trigonometry.

Consider first the relationship between the sine, cosine and tangent of any angle. With P(x,y)P(x, y) as in Figure 5.29,

sin⁡θ=yOP,cos⁡θ=xOP,tan⁡θ=yx.\sin\theta = \frac{y}{OP}, \qquad \cos\theta = \frac{x}{OP}, \qquad \tan\theta = \frac{y}{x} .

But yx=yOP÷xOP\dfrac{y}{x} = \dfrac{y}{OP} \div \dfrac{x}{OP}, so for all angles

tan⁡θ=sin⁡θcos⁡θ,and similarlycot⁡θ=cos⁡θsin⁡θ.\tan\theta = \frac{\sin\theta}{\cos\theta}, \qquad\text{and similarly}\qquad \cot\theta = \frac{\cos\theta}{\sin\theta} .
xyP(x, y)QOxyOPθ
Figure 5.29. The right-angled triangle OPQOPQ for any position of PP.

The Pythagorean Identities

For any position of OPOP a right-angled triangle OPQOPQ can be drawn, for which, by Pythagoras,

x2+y2=OP2.x^2 + y^2 = OP^2 .

Dividing throughout, in turn, by OP2OP^2, x2x^2 and y2y^2 gives

(xOP)2+(yOP)2=1,1+(yx)2=(OPx)2,(xy)2+1=(OPy)2.\left(\frac{x}{OP}\right)^2 + \left(\frac{y}{OP}\right)^2 = 1, \qquad 1 + \left(\frac{y}{x}\right)^2 = \left(\frac{OP}{x}\right)^2, \qquad \left(\frac{x}{y}\right)^2 + 1 = \left(\frac{OP}{y}\right)^2 .

The use of brackets when raising a trigonometric ratio to a power can be avoided by writing cos⁡2θ\cos^2\theta for (cos⁡θ)2(\cos\theta)^2, and so on, and then these become

cos⁡2θ+sin⁡2θ=1,1+tan⁡2θ=sec⁡2θ,cot⁡2θ+1=cosec⁡2θ.\cos^2\theta + \sin^2\theta = 1, \qquad 1 + \tan^2\theta = \sec^2\theta, \qquad \cot^2\theta + 1 = \operatorname{cosec}^2\theta .

These relationships are valid for any position of OPOP, that is for all angles. They are very useful in the solution of certain trigonometric equations.

Example 5.34.

Solve the equation 2cos⁡2θ−sin⁡θ=12\cos^2\theta - \sin\theta = 1 for values of θ\theta between 00 and 2π2\pi.

Using cos⁡2θ+sin⁡2θ=1\cos^2\theta + \sin^2\theta = 1 gives

2(1−sin⁡2θ)−sin⁡θ=1,or2sin⁡2θ+sin⁡θ−1=0.2(1 - \sin^2\theta) - \sin\theta = 1, \qquad\text{or}\qquad 2\sin^2\theta + \sin\theta - 1 = 0 .

This is now a quadratic equation in sin⁡θ\sin\theta, of the form 2x2+x−1=02x^2 + x - 1 = 0, hence

(2sin⁡θ−1)(sin⁡θ+1)=0,sin⁡θ=12 or sin⁡θ=−1.(2\sin\theta - 1)(\sin\theta + 1) = 0, \qquad \sin\theta = \tfrac12 \ \text{or}\ \sin\theta = -1 .

If sin⁡θ=12\sin\theta = \frac12, θ=π6,5π6\theta = \dfrac{\pi}{6}, \dfrac{5\pi}{6}, and if sin⁡θ=−1\sin\theta = -1, θ=3π2\theta = \dfrac{3\pi}{2}. Therefore the solution of the equation is θ=π6,5π6,3π2\theta = \dfrac{\pi}{6}, \dfrac{5\pi}{6}, \dfrac{3\pi}{2}.

Example 5.35.

If 3sec⁡2θ−5tan⁡θ−4=03\sec^2\theta - 5\tan\theta - 4 = 0, find the general solution.

Using 1+tan⁡2θ=sec⁡2θ1 + \tan^2\theta = \sec^2\theta we have

3(1+tan⁡2θ)−5tan⁡θ−4=0,that is3tan⁡2θ−5tan⁡θ−1=0.3(1 + \tan^2\theta) - 5\tan\theta - 4 = 0, \qquad\text{that is}\qquad 3\tan^2\theta - 5\tan\theta - 1 = 0 .

Again we have a quadratic equation, but because it has no simple factors we solve it by the formula:

tan⁡θ=5±25+126=1.8471 or −0.1805.\tan\theta = \frac{5 \pm \sqrt{25 + 12}}{6} = 1.8471 \ \text{or}\ -0.1805 .

If tan⁡θ=1.8471\tan\theta = 1.8471 the principal solution is θ=61.57∘\theta = 61.57^\circ, and if tan⁡θ=−0.1805\tan\theta = -0.1805 it is θ=−10.23∘\theta = -10.23^\circ. The complete general solution is therefore

θ=180n∘+61.57∘orθ=180n∘−10.23∘.\theta = 180n^\circ + 61.57^\circ \qquad\text{or}\qquad \theta = 180n^\circ - 10.23^\circ .

Other applications of the standard identities include the derivation of further trigonometric relationships, the elimination of trigonometric terms from pairs of equations, and the calculation of the remaining trigonometric ratios of an angle for which only one ratio is known.

Example 5.36.

Show that (1−cos⁡A)(1+sec⁡A)=sin⁡Atan⁡A(1 - \cos A)(1 + \sec A) = \sin A\tan A.

Because this relationship has yet to be established, we must not assume that it is true by using the complete identity in our working: the left- and right-hand sides must be kept apart throughout. Considering the left-hand side,

(1−cos⁡A)(1+sec⁡A)=1+sec⁡A−cos⁡A−cos⁡Asec⁡A=1+sec⁡A−cos⁡A−1=1cos⁡A−cos⁡A=1−cos⁡2Acos⁡A=sin⁡2Acos⁡A=sin⁡Atan⁡A,\begin{aligned} (1 - \cos A)(1 + \sec A) &= 1 + \sec A - \cos A - \cos A\sec A = 1 + \sec A - \cos A - 1 \\ &= \frac{1}{\cos A} - \cos A = \frac{1 - \cos^2 A}{\cos A} = \frac{\sin^2 A}{\cos A} = \sin A\tan A, \end{aligned}

which is identical to the right-hand side.

Example 5.37.

Show that (cosec⁡A−sin⁡A)(sec⁡A−cos⁡A)=1tan⁡A+cot⁡A(\operatorname{cosec} A - \sin A)(\sec A - \cos A) = \dfrac{1}{\tan A + \cot A}.

Considering the left-hand side,

(cosec⁡A−sin⁡A)(sec⁡A−cos⁡A)=(1−sin⁡2Asin⁡A)(1−cos⁡2Acos⁡A)=cos⁡2Asin⁡A⋅sin⁡2Acos⁡A=cos⁡Asin⁡A.(\operatorname{cosec} A - \sin A)(\sec A - \cos A) = \left(\frac{1 - \sin^2 A}{\sin A}\right)\left(\frac{1 - \cos^2 A}{\cos A}\right) = \frac{\cos^2 A}{\sin A} \cdot \frac{\sin^2 A}{\cos A} = \cos A\sin A .

This is already a very simple form, but it is not obviously identical to the right-hand side, so this time we work independently on the right-hand side:

1tan⁡A+cot⁡A=1÷(sin⁡Acos⁡A+cos⁡Asin⁡A)=1÷sin⁡2A+cos⁡2Acos⁡Asin⁡A=cos⁡Asin⁡A.\frac{1}{\tan A + \cot A} = 1 \div \left(\frac{\sin A}{\cos A} + \frac{\cos A}{\sin A}\right) = 1 \div \frac{\sin^2 A + \cos^2 A}{\cos A\sin A} = \cos A\sin A .

Since both sides reduce to cos⁡Asin⁡A\cos A\sin A, they are identical.

Example 5.38.

Eliminate θ\theta from the equations x=2cos⁡θx = 2\cos\theta and y=3sin⁡θy = 3\sin\theta.

Here cos⁡θ=x2\cos\theta = \dfrac{x}{2} and sin⁡θ=y3\sin\theta = \dfrac{y}{3}. Using cos⁡2θ+sin⁡2θ=1\cos^2\theta + \sin^2\theta = 1 gives

x24+y29=1,that is9x2+4y2=36.\frac{x^2}{4} + \frac{y^2}{9} = 1, \qquad\text{that is}\qquad 9x^2 + 4y^2 = 36 .

In this problem both xx and yy depend on θ\theta, a variable angle. Used in this way, θ\theta is called a parameter.

Example 5.39.

If sin⁡A=13\sin A = \frac13 and AA is obtuse, find cos⁡A\cos A and cot⁡A\cot A without using a calculator.

Using cos⁡2A+sin⁡2A=1\cos^2 A + \sin^2 A = 1 gives cos⁡2A+19=1\cos^2 A + \frac19 = 1, so cos⁡2A=89\cos^2 A = \frac89 and cos⁡A=±223\cos A = \pm\dfrac{2\sqrt2}{3}. But AA is obtuse, so cos⁡A\cos A is negative. Hence

cos⁡A=−223,cot⁡A=cos⁡Asin⁡A=−22.\cos A = -\frac{2\sqrt2}{3}, \qquad \cot A = \frac{\cos A}{\sin A} = -2\sqrt2 .

This type of problem can often be done more directly by drawing the appropriate right-angled triangle and using Pythagoras.

Problem 5.12.

  1. Solve the equation tan⁡θ+cot⁡θ=2\tan\theta + \cot\theta = 2 for angles in the range −180∘⩽θ⩽180∘-180^\circ \leqslant \theta \leqslant 180^\circ.
  2. Find the general solution of the equation 5cos⁡θ−4sin⁡2θ=25\cos\theta - 4\sin^2\theta = 2.
  3. Show that cot⁡θ+tan⁡θ=sec⁡θcosec⁡θ\cot\theta + \tan\theta = \sec\theta\operatorname{cosec}\theta, and that sin⁡A1+cos⁡A=1−cos⁡Asin⁡A\dfrac{\sin A}{1 + \cos A} = \dfrac{1 - \cos A}{\sin A}.
  4. Eliminate θ\theta from the equations x=4sec⁡θx = 4\sec\theta and y=5tan⁡θy = 5\tan\theta.

Compound Angle Identities

It is often useful to be able to express the trigonometric ratios of angles such as A+BA + B or A−BA - B in terms of the ratios of AA and of BB. At first sight it is dangerously easy to think, for instance, that sin⁡(A+B)\sin(A + B) is sin⁡A+sin⁡B\sin A + \sin B. That this is false can be seen by considering

sin⁡(45∘+45∘)=sin⁡90∘=1,whereassin⁡45∘+sin⁡45∘=22+22=2≠1.\sin(45^\circ + 45^\circ) = \sin 90^\circ = 1, \qquad\text{whereas}\qquad \sin 45^\circ + \sin 45^\circ = \frac{\sqrt2}{2} + \frac{\sqrt2}{2} = \sqrt2 \neq 1 .

So the sine function is not distributive, and similarly for the other trigonometric ratios. The correct expression is

sin⁡(A+B)=sin⁡Acos⁡B+cos⁡Asin⁡B.\sin(A + B) = \sin A\cos B + \cos A\sin B .

We show this geometrically when AA and BB are both acute, using Figure 5.30. The right-angled triangles OPQOPQ and OQROQR contain the angles AA and BB, RTRT is perpendicular to OPOP, and QSQS is perpendicular to RTRT. Since QSQS is parallel to OPOP, the angle SQOSQO is AA, so the angle SQRSQR is 90∘−A90^\circ - A and the angle SRQSRQ is equal to AA. Since TS=PQTS = PQ,

sin⁡(A+B)=TROR=TS+SROR=PQOR+SROR=PQOQ⋅OQOR+SRQR⋅QROR=sin⁡Acos⁡B+cos⁡Asin⁡B.\sin(A + B) = \frac{TR}{OR} = \frac{TS + SR}{OR} = \frac{PQ}{OR} + \frac{SR}{OR} = \frac{PQ}{OQ} \cdot \frac{OQ}{OR} + \frac{SR}{QR} \cdot \frac{QR}{OR} = \sin A\cos B + \cos A\sin B .
ABAOPQRTS
Figure 5.30. The construction for sin⁡(A+B)\sin(A + B) when AA and BB are acute.

Accepting at this stage that the formula is valid for all angles, it can be adapted to give the full set of compound angle identities.

  1. Replacing BB by −B-B, and using cos⁡(−B)=cos⁡B\cos(-B) = \cos B and sin⁡(−B)=−sin⁡B\sin(-B) = -\sin B, gives sin⁡(A−B)=sin⁡Acos⁡B−cos⁡Asin⁡B\sin(A - B) = \sin A\cos B - \cos A\sin B.
  2. Replacing AA by π2−A\dfrac{\pi}{2} - A in this identity, and using the complementary ratios, gives sin⁡(π2−(A+B))=cos⁡Acos⁡B−sin⁡Asin⁡B\sin\left(\dfrac{\pi}{2} - (A + B)\right) = \cos A\cos B - \sin A\sin B, that is cos⁡(A+B)=cos⁡Acos⁡B−sin⁡Asin⁡B\cos(A + B) = \cos A\cos B - \sin A\sin B.
  3. Replacing BB by −B-B in this identity gives cos⁡(A−B)=cos⁡Acos⁡B+sin⁡Asin⁡B\cos(A - B) = \cos A\cos B + \sin A\sin B.
  4. Dividing sin⁡(A+B)\sin(A + B) by cos⁡(A+B)\cos(A + B), and then dividing the numerator and the denominator by cos⁡Acos⁡B\cos A\cos B, gives tan⁡(A+B)=tan⁡A+tan⁡B1−tan⁡Atan⁡B\tan(A + B) = \dfrac{\tan A + \tan B}{1 - \tan A\tan B}.
  5. Replacing BB by −B-B in this identity, with tan⁡(−B)=−tan⁡B\tan(-B) = -\tan B, gives tan⁡(A−B)=tan⁡A−tan⁡B1+tan⁡Atan⁡B\tan(A - B) = \dfrac{\tan A - \tan B}{1 + \tan A\tan B}.

Collating these results we have

sin⁡(A+B)=sin⁡Acos⁡B+cos⁡Asin⁡B,sin⁡(A−B)=sin⁡Acos⁡B−cos⁡Asin⁡B,cos⁡(A+B)=cos⁡Acos⁡B−sin⁡Asin⁡B,cos⁡(A−B)=cos⁡Acos⁡B+sin⁡Asin⁡B,tan⁡(A+B)=tan⁡A+tan⁡B1−tan⁡Atan⁡B,tan⁡(A−B)=tan⁡A−tan⁡B1+tan⁡Atan⁡B.\begin{aligned} \sin(A + B) &= \sin A\cos B + \cos A\sin B, & \sin(A - B) &= \sin A\cos B - \cos A\sin B, \\ \cos(A + B) &= \cos A\cos B - \sin A\sin B, & \cos(A - B) &= \cos A\cos B + \sin A\sin B, \\ \tan(A + B) &= \frac{\tan A + \tan B}{1 - \tan A\tan B}, & \tan(A - B) &= \frac{\tan A - \tan B}{1 + \tan A\tan B} . \end{aligned}

The similarity between the pairs of identities for A+BA + B and A−BA - B makes it clear that care must be taken with signs when using these formulae.

Example 5.40.

Without using a calculator, evaluate sin⁡75∘\sin 75^\circ, cos⁡105∘\cos 105^\circ and tan⁡(−15∘)\tan(-15^\circ).

sin⁡75∘=sin⁡(45∘+30∘)=sin⁡45∘cos⁡30∘+cos⁡45∘sin⁡30∘=12⋅32+12⋅12=3+122,cos⁡105∘=cos⁡(60∘+45∘)=cos⁡60∘cos⁡45∘−sin⁡60∘sin⁡45∘=12⋅12−32⋅12=1−322,tan⁡(−15∘)=tan⁡(45∘−60∘)=tan⁡45∘−tan⁡60∘1+tan⁡45∘tan⁡60∘=1−31+3=3−2,\begin{aligned} \sin 75^\circ &= \sin(45^\circ + 30^\circ) = \sin 45^\circ\cos 30^\circ + \cos 45^\circ\sin 30^\circ = \frac{1}{\sqrt2} \cdot \frac{\sqrt3}{2} + \frac{1}{\sqrt2} \cdot \frac12 = \frac{\sqrt3 + 1}{2\sqrt2}, \\ \cos 105^\circ &= \cos(60^\circ + 45^\circ) = \cos 60^\circ\cos 45^\circ - \sin 60^\circ\sin 45^\circ = \frac12 \cdot \frac{1}{\sqrt2} - \frac{\sqrt3}{2} \cdot \frac{1}{\sqrt2} = \frac{1 - \sqrt3}{2\sqrt2}, \\ \tan(-15^\circ) &= \tan(45^\circ - 60^\circ) = \frac{\tan 45^\circ - \tan 60^\circ}{1 + \tan 45^\circ\tan 60^\circ} = \frac{1 - \sqrt3}{1 + \sqrt3} = \sqrt3 - 2, \end{aligned}

the last after rationalising the denominator. The value of cos⁡105∘\cos 105^\circ is negative, which is consistent with the cosine of an angle in the second quadrant. In each part there are alternative compound angles which could be used, such as 75∘=120∘−45∘75^\circ = 120^\circ - 45^\circ, 105∘=150∘−45∘105^\circ = 150^\circ - 45^\circ and −15∘=30∘−45∘-15^\circ = 30^\circ - 45^\circ.

Example 5.41.

AA is obtuse and sin⁡A=35\sin A = \frac35, and BB is acute and sin⁡B=1213\sin B = \frac{12}{13}. Without finding the values of AA and BB, evaluate cos⁡(A+B)\cos(A + B) and tan⁡(A−B)\tan(A - B).

In order to use the compound angle formulae we need cos⁡A\cos A, cos⁡B\cos B, tan⁡A\tan A and tan⁡B\tan B. These are most simply obtained by using Pythagoras in the appropriate right-angled triangles, with sides 33, 44, 55 and 55, 1212, 1313, remembering that AA is obtuse:

cos⁡A=−45,tan⁡A=−34,cos⁡B=513,tan⁡B=125.\cos A = -\frac45, \quad \tan A = -\frac34, \qquad \cos B = \frac{5}{13}, \quad \tan B = \frac{12}{5} .

Then

cos⁡(A+B)=cos⁡Acos⁡B−sin⁡Asin⁡B=(−45)(513)−(35)(1213)=−5665,tan⁡(A−B)=tan⁡A−tan⁡B1+tan⁡Atan⁡B=−34−1251+(−34)(125)=6316.\begin{aligned} \cos(A + B) &= \cos A\cos B - \sin A\sin B = \left(-\frac45\right)\left(\frac{5}{13}\right) - \left(\frac35\right)\left(\frac{12}{13}\right) = -\frac{56}{65}, \\ \tan(A - B) &= \frac{\tan A - \tan B}{1 + \tan A\tan B} = \frac{-\frac34 - \frac{12}{5}}{1 + \left(-\frac34\right)\left(\frac{12}{5}\right)} = \frac{63}{16} . \end{aligned}

Example 5.42.

Show that

sin⁡(A−B)cos⁡Acos⁡B+sin⁡(B−C)cos⁡Bcos⁡C+sin⁡(C−A)cos⁡Ccos⁡A=0.\frac{\sin(A - B)}{\cos A\cos B} + \frac{\sin(B - C)}{\cos B\cos C} + \frac{\sin(C - A)}{\cos C\cos A} = 0 .

Expanding each numerator, the left-hand side becomes

sin⁡Acos⁡B−cos⁡Asin⁡Bcos⁡Acos⁡B+sin⁡Bcos⁡C−cos⁡Bsin⁡Ccos⁡Bcos⁡C+sin⁡Ccos⁡A−cos⁡Csin⁡Acos⁡Ccos⁡A,\frac{\sin A\cos B - \cos A\sin B}{\cos A\cos B} + \frac{\sin B\cos C - \cos B\sin C}{\cos B\cos C} + \frac{\sin C\cos A - \cos C\sin A}{\cos C\cos A},

which is

(tan⁡A−tan⁡B)+(tan⁡B−tan⁡C)+(tan⁡C−tan⁡A)=0.(\tan A - \tan B) + (\tan B - \tan C) + (\tan C - \tan A) = 0 .

Example 5.43.

Solve the equation 2cos⁡θ=sin⁡(θ+30∘)2\cos\theta = \sin(\theta + 30^\circ), giving the general values of θ\theta.

It is very important to appreciate that this is not an identity: only certain values of θ\theta satisfy the equation.

2cos⁡θ=sin⁡θcos⁡30∘+cos⁡θsin⁡30∘=32sin⁡θ+12cos⁡θ,2\cos\theta = \sin\theta\cos 30^\circ + \cos\theta\sin 30^\circ = \frac{\sqrt3}{2}\sin\theta + \frac12\cos\theta,

therefore

32cos⁡θ=32sin⁡θ,sotan⁡θ=3.\frac32\cos\theta = \frac{\sqrt3}{2}\sin\theta, \qquad\text{so}\qquad \tan\theta = \sqrt3 .

The principal solution is θ=60∘\theta = 60^\circ, so the general solution is θ=180n∘+60∘\theta = 180n^\circ + 60^\circ.

Problem 5.13.

  1. Without using a calculator, evaluate cos⁡80∘cos⁡20∘+sin⁡80∘sin⁡20∘\cos 80^\circ\cos 20^\circ + \sin 80^\circ\sin 20^\circ and sin⁡165∘\sin 165^\circ.
  2. Show that sin⁡(A+B)cos⁡Acos⁡B=tan⁡A+tan⁡B\dfrac{\sin(A + B)}{\cos A\cos B} = \tan A + \tan B.
  3. Solve the equation sin⁡(x+60∘)=cos⁡x\sin(x + 60^\circ) = \cos x, giving the angles from 0∘0^\circ to 360∘360^\circ.

Double Angle Identities

The compound angle formulae deal with any two angles AA and BB, so they can be used for two equal angles. Replacing BB by AA in the formulae for A+BA + B gives

sin⁡2A=2sin⁡Acos⁡A,cos⁡2A=cos⁡2A−sin⁡2A,tan⁡2A=2tan⁡A1−tan⁡2A.\sin 2A = 2\sin A\cos A, \qquad \cos 2A = \cos^2 A - \sin^2 A, \qquad \tan 2A = \frac{2\tan A}{1 - \tan^2 A} .

The second of these can be expressed in several forms, because

cos⁡2A−sin⁡2A=(1−sin⁡2A)−sin⁡2A=1−2sin⁡2A,cos⁡2A−sin⁡2A=cos⁡2A−(1−cos⁡2A)=2cos⁡2A−1.\cos^2 A - \sin^2 A = (1 - \sin^2 A) - \sin^2 A = 1 - 2\sin^2 A, \qquad \cos^2 A - \sin^2 A = \cos^2 A - (1 - \cos^2 A) = 2\cos^2 A - 1 .

Thus

cos⁡2A=cos⁡2A−sin⁡2A=1−2sin⁡2A=2cos⁡2A−1,\cos 2A = \cos^2 A - \sin^2 A = 1 - 2\sin^2 A = 2\cos^2 A - 1,

and these alternative expressions can themselves be rearranged to give

2sin⁡2A=1−cos⁡2A,2cos⁡2A=1+cos⁡2A.2\sin^2 A = 1 - \cos 2A, \qquad 2\cos^2 A = 1 + \cos 2A .

Complete familiarity with all the double angle formulae, including all the alternative forms of cos⁡2A\cos 2A, is essential. They are probably the most useful of all the trigonometric identities for simplifying trigonometric functions.

Example 5.44.

Find the general solution of the equation cos⁡2x+3sin⁡x=2\cos 2x + 3\sin x = 2.

Using cos⁡2x=1−2sin⁡2x\cos 2x = 1 - 2\sin^2 x gives

1−2sin⁡2x+3sin⁡x=2,2sin⁡2x−3sin⁡x+1=0,(2sin⁡x−1)(sin⁡x−1)=0,1 - 2\sin^2 x + 3\sin x = 2, \qquad 2\sin^2 x - 3\sin x + 1 = 0, \qquad (2\sin x - 1)(\sin x - 1) = 0,

so sin⁡x=12\sin x = \frac12 or sin⁡x=1\sin x = 1.

  1. For sin⁡x=12\sin x = \frac12 the principal solution is x=π6x = \dfrac{\pi}{6}, and the general solution is x=2nπ+π6x = 2n\pi + \dfrac{\pi}{6} or x=(2n+1)π−π6x = (2n + 1)\pi - \dfrac{\pi}{6}.
  2. For sin⁡x=1\sin x = 1 the principal solution is x=π2x = \dfrac{\pi}{2}, and the general solution is x=2nπ+π2x = 2n\pi + \dfrac{\pi}{2}.

The full solution is therefore x=2nπ+π6x = 2n\pi + \dfrac{\pi}{6}, (2n+1)π−π6(2n + 1)\pi - \dfrac{\pi}{6} or 2nπ+π22n\pi + \dfrac{\pi}{2}.

Example 5.45.

If tan⁡θ=34\tan\theta = \frac34 and θ\theta is acute, find the values of tan⁡2θ\tan 2\theta, tan⁡4θ\tan 4\theta and tan⁡θ2\tan\dfrac{\theta}{2}.

In this problem we use tan⁡2A=2tan⁡A1−tan⁡2A\tan 2A = \dfrac{2\tan A}{1 - \tan^2 A} for three cases.

A=θ:tan⁡2θ=2×341−916=247,A=2θ:tan⁡4θ=2×2471−57649=−336527.\begin{aligned} A = \theta: &\qquad \tan 2\theta = \frac{2 \times \frac34}{1 - \frac{9}{16}} = \frac{24}{7}, \\ A = 2\theta: &\qquad \tan 4\theta = \frac{2 \times \frac{24}{7}}{1 - \frac{576}{49}} = -\frac{336}{527} . \end{aligned}

For A=θ2A = \dfrac{\theta}{2}, writing t=tan⁡θ2t = \tan\dfrac{\theta}{2},

2t1−t2=34,8t=3−3t2,3t2+8t−3=0,(3t−1)(t+3)=0,\frac{2t}{1 - t^2} = \frac34, \qquad 8t = 3 - 3t^2, \qquad 3t^2 + 8t - 3 = 0, \qquad (3t - 1)(t + 3) = 0,

so t=13t = \frac13 or t=−3t = -3. But θ\theta is acute, so θ2\dfrac{\theta}{2} is acute and its tangent is positive. Therefore tan⁡θ2=13\tan\dfrac{\theta}{2} = \dfrac13.

Example 5.46.

Show that sin⁡3A=3sin⁡A−4sin⁡3A\sin 3A = 3\sin A - 4\sin^3 A.

sin⁡3A=sin⁡(2A+A)=sin⁡2Acos⁡A+cos⁡2Asin⁡A=2sin⁡Acos⁡2A+(1−2sin⁡2A)sin⁡A=2sin⁡A(1−sin⁡2A)+sin⁡A−2sin⁡3A=3sin⁡A−4sin⁡3A.\begin{aligned} \sin 3A = \sin(2A + A) &= \sin 2A\cos A + \cos 2A\sin A = 2\sin A\cos^2 A + (1 - 2\sin^2 A)\sin A \\ &= 2\sin A(1 - \sin^2 A) + \sin A - 2\sin^3 A = 3\sin A - 4\sin^3 A . \end{aligned}

This is quite a useful identity, and it is worth remembering.

Example 5.47.

Eliminate θ\theta from the equations x=cos⁡2θx = \cos 2\theta and y=sec⁡θy = \sec\theta.

Using cos⁡2θ=2cos⁡2θ−1\cos 2\theta = 2\cos^2\theta - 1, and cos⁡θ=1y\cos\theta = \dfrac{1}{y},

x=2y2−1,hence(x+1)y2=2.x = \frac{2}{y^2} - 1, \qquad\text{hence}\qquad (x + 1)y^2 = 2 .

This is a Cartesian equation, obtained by eliminating the parameter θ\theta from a pair of parametric equations.

Problem 5.14.

  1. Express as a single trigonometric ratio 2sin⁡14∘cos⁡14∘2\sin 14^\circ\cos 14^\circ, 1−2sin⁡240∘1 - 2\sin^2 40^\circ and 1+tan⁡x1−tan⁡x\dfrac{1 + \tan x}{1 - \tan x}.
  2. Solve the equation cos⁡2x=sin⁡x\cos 2x = \sin x for angles in the range 0∘⩽x⩽360∘0^\circ \leqslant x \leqslant 360^\circ, and state the general solution.
  3. Show that cos⁡3θ=4cos⁡3θ−3cos⁡θ\cos 3\theta = 4\cos^3\theta - 3\cos\theta.

Identities and the Inverse Functions

The double angle and compound angle identities are often useful in simplifying expressions, or solving equations, which contain inverse trigonometric functions.

Example 5.48.

Find xx if arcsin⁡x+arccos⁡x2=5π6\arcsin x + \arccos\dfrac{x}{2} = \dfrac{5\pi}{6}.

Let arcsin⁡x=θ\arcsin x = \theta, so that sin⁡θ=x\sin\theta = x, and cos⁡θ=1−x2\cos\theta = \sqrt{1 - x^2} since −π2⩽θ⩽π2-\dfrac{\pi}{2} \leqslant \theta \leqslant \dfrac{\pi}{2}. Let arccos⁡x2=φ\arccos\dfrac{x}{2} = \varphi, so that cos⁡φ=x2\cos\varphi = \dfrac{x}{2}, and sin⁡φ=4−x22\sin\varphi = \dfrac{\sqrt{4 - x^2}}{2} since 0⩽φ⩽π0 \leqslant \varphi \leqslant \pi. The given equation then becomes θ+φ=5π6\theta + \varphi = \dfrac{5\pi}{6}, so

sin⁡(θ+φ)=12,sin⁡θcos⁡φ+cos⁡θsin⁡φ=12,\sin(\theta + \varphi) = \frac12, \qquad \sin\theta\cos\varphi + \cos\theta\sin\varphi = \frac12,

that is

x22+1−x2 4−x22=12,1−x2 4−x2=1−x2.\frac{x^2}{2} + \frac{\sqrt{1 - x^2}\,\sqrt{4 - x^2}}{2} = \frac12, \qquad \sqrt{1 - x^2}\,\sqrt{4 - x^2} = 1 - x^2 .

Squaring both sides, and not cancelling 1−x2\sqrt{1 - x^2}, which would lose solutions, gives

(1−x2)(4−x2)=(1−x2)2,(1−x2)[(4−x2)−(1−x2)]=3(1−x2)=0,(1 - x^2)(4 - x^2) = (1 - x^2)^2, \qquad (1 - x^2)\bigl[(4 - x^2) - (1 - x^2)\bigr] = 3(1 - x^2) = 0,

so x=±1x = \pm 1. But arcsin⁡(−1)=−π2\arcsin(-1) = -\dfrac{\pi}{2} and arccos⁡(−12)=2π3\arccos\left(-\dfrac12\right) = \dfrac{2\pi}{3}, whose sum is π6\dfrac{\pi}{6}, so x=−1x = -1 does not satisfy the given equation. The only solution is x=1x = 1.

Example 5.49.

Show that arctan⁡3+2arctan⁡2=π+arccot⁡3\arctan 3 + 2\arctan 2 = \pi + \operatorname{arccot} 3.

Let θ=arctan⁡3\theta = \arctan 3, so that tan⁡θ=3\tan\theta = 3 and π4<θ<π2\dfrac{\pi}{4} < \theta < \dfrac{\pi}{2}, and let φ=arctan⁡2\varphi = \arctan 2, so that tan⁡φ=2\tan\varphi = 2 and π4<φ<π2\dfrac{\pi}{4} < \varphi < \dfrac{\pi}{2}. The left-hand side is then θ+2φ\theta + 2\varphi. Now

tan⁡2φ=2×21−4=−43,tan⁡(θ+2φ)=tan⁡θ+tan⁡2φ1−tan⁡θtan⁡2φ=3−431−3(−43)=535=13.\tan 2\varphi = \frac{2 \times 2}{1 - 4} = -\frac43, \qquad \tan(\theta + 2\varphi) = \frac{\tan\theta + \tan 2\varphi}{1 - \tan\theta\tan 2\varphi} = \frac{3 - \frac43}{1 - 3\left(-\frac43\right)} = \frac{\frac53}{5} = \frac13 .

But π4<θ<π2\dfrac{\pi}{4} < \theta < \dfrac{\pi}{2} and π4<φ<π2\dfrac{\pi}{4} < \varphi < \dfrac{\pi}{2} give 3π4<θ+2φ<3π2\dfrac{3\pi}{4} < \theta + 2\varphi < \dfrac{3\pi}{2}, and the only angle in this range with tangent 13\frac13 is π+arctan⁡13\pi + \arctan\frac13. So

arctan⁡3+2arctan⁡2=π+arctan⁡13=π+arccot⁡3.\arctan 3 + 2\arctan 2 = \pi + \arctan\frac13 = \pi + \operatorname{arccot} 3 .

Example 5.50.

Simplify arctan⁡x+arctan⁡1−x1+x\arctan x + \arctan\dfrac{1 - x}{1 + x}, where x>−1x > -1.

Let α=arctan⁡x\alpha = \arctan x and β=arctan⁡1−x1+x\beta = \arctan\dfrac{1 - x}{1 + x}, so that tan⁡α=x\tan\alpha = x and tan⁡β=1−x1+x\tan\beta = \dfrac{1 - x}{1 + x}, and we have to simplify α+β\alpha + \beta. Using tan⁡(α+β)=tan⁡α+tan⁡β1−tan⁡αtan⁡β\tan(\alpha + \beta) = \dfrac{\tan\alpha + \tan\beta}{1 - \tan\alpha\tan\beta} gives

tan⁡(α+β)=x+1−x1+x1−x⋅1−x1+x=x+x2+1−x1+x−x+x2=1.\tan(\alpha + \beta) = \frac{x + \dfrac{1 - x}{1 + x}}{1 - x \cdot \dfrac{1 - x}{1 + x}} = \frac{x + x^2 + 1 - x}{1 + x - x + x^2} = 1 .

Since x>−1x > -1, both xx and 1−x1+x\dfrac{1 - x}{1 + x} are greater than −1-1, so α\alpha and β\beta each lie between −π4-\dfrac{\pi}{4} and π2\dfrac{\pi}{2}, and α+β\alpha + \beta lies between −π2-\dfrac{\pi}{2} and π\pi. The only angle in that range whose tangent is 11 is π4\dfrac{\pi}{4}. Thus

arctan⁡x+arctan⁡1−x1+x=π4.\arctan x + \arctan\frac{1 - x}{1 + x} = \frac{\pi}{4} .

Problem 5.15.

  1. Show that arctan⁡13+arctan⁡12=π4\arctan\dfrac13 + \arctan\dfrac12 = \dfrac{\pi}{4}.
  2. Solve the equation arctan⁡(1+x)+arctan⁡(1−x)=arctan⁡2\arctan(1 + x) + \arctan(1 - x) = \arctan 2.
  3. Simplify sin⁡(2arctan⁡x)\sin(2\arctan x).

The Half Angle Identities

We already know that tan⁡2A=2tan⁡A1−tan⁡2A\tan 2A = \dfrac{2\tan A}{1 - \tan^2 A}. Dividing sin⁡2A=2sin⁡Acos⁡A\sin 2A = 2\sin A\cos A and cos⁡2A=cos⁡2A−sin⁡2A\cos 2A = \cos^2 A - \sin^2 A by cos⁡2A+sin⁡2A\cos^2 A + \sin^2 A, which is 11, and then dividing the numerator and the denominator by cos⁡2A\cos^2 A, gives also

sin⁡2A=2tan⁡A1+tan⁡2A,cos⁡2A=1−tan⁡2A1+tan⁡2A.\sin 2A = \frac{2\tan A}{1 + \tan^2 A}, \qquad \cos 2A = \frac{1 - \tan^2 A}{1 + \tan^2 A} .

If we replace 2A2A by θ\theta and use tt to denote tan⁡θ2\tan\dfrac{\theta}{2}, we have

tan⁡θ=2t1−t2,sin⁡θ=2t1+t2,cos⁡θ=1−t21+t2.\tan\theta = \frac{2t}{1 - t^2}, \qquad \sin\theta = \frac{2t}{1 + t^2}, \qquad \cos\theta = \frac{1 - t^2}{1 + t^2} .

These three identities allow all the trigonometric ratios of any one angle to be expressed in terms of a common variable tt. In problems where none of the identities used so far can be applied, this group can be helpful.

Example 5.51.

Solve the equation sin⁡θ+2cos⁡θ=1\sin\theta + 2\cos\theta = 1 for angles between 0∘0^\circ and 360∘360^\circ.

With t=tan⁡θ2t = \tan\dfrac{\theta}{2},

2t1+t2+2(1−t2)1+t2=1,2t+2−2t2=1+t2,3t2−2t−1=0,(3t+1)(t−1)=0.\frac{2t}{1 + t^2} + \frac{2(1 - t^2)}{1 + t^2} = 1, \qquad 2t + 2 - 2t^2 = 1 + t^2, \qquad 3t^2 - 2t - 1 = 0, \qquad (3t + 1)(t - 1) = 0 .

Therefore either tan⁡θ2=−13\tan\dfrac{\theta}{2} = -\dfrac13 or tan⁡θ2=1\tan\dfrac{\theta}{2} = 1. The range of values specified for θ\theta is 0∘0^\circ to 360∘360^\circ, so the range of values required for θ2\dfrac{\theta}{2} is 0∘0^\circ to 180∘180^\circ. Within this range tan⁡θ2=−13\tan\dfrac{\theta}{2} = -\dfrac13 gives θ2=161.57∘\dfrac{\theta}{2} = 161.57^\circ, and tan⁡θ2=1\tan\dfrac{\theta}{2} = 1 gives θ2=45∘\dfrac{\theta}{2} = 45^\circ. Thus θ=323.13∘,90∘\theta = 323.13^\circ, 90^\circ.

Extra care is sometimes needed with this method, as a very similar equation shows. Consider

sin⁡θ−cos⁡θ=1.\sin\theta - \cos\theta = 1 .

Using t=tan⁡θ2t = \tan\dfrac{\theta}{2} gives

2t1+t2−1−t21+t2=1,2t−1+t2=1+t2.\frac{2t}{1 + t^2} - \frac{1 - t^2}{1 + t^2} = 1, \qquad 2t - 1 + t^2 = 1 + t^2 .

The two t2t^2 terms cancel, leaving 2t=22t = 2, so t=1t = 1. But t=tan⁡θ2t = \tan\dfrac{\theta}{2} is not defined when θ2\dfrac{\theta}{2} is an odd multiple of π2\dfrac{\pi}{2}, that is when θ=(2n+1)π\theta = (2n + 1)\pi, and these angles have to be checked separately: sin⁡π−cos⁡π=0+1=1\sin\pi - \cos\pi = 0 + 1 = 1, so they are solutions as well. Hence tan⁡θ2=1\tan\dfrac{\theta}{2} = 1, or tan⁡θ2\tan\dfrac{\theta}{2} is undefined, giving

θ2=nπ+π4orθ2=nπ+π2,that isθ=2nπ+π2orθ=(2n+1)π.\frac{\theta}{2} = n\pi + \frac{\pi}{4} \quad\text{or}\quad \frac{\theta}{2} = n\pi + \frac{\pi}{2}, \qquad\text{that is}\qquad \theta = 2n\pi + \frac{\pi}{2} \quad\text{or}\quad \theta = (2n + 1)\pi .

Remark.

When the half angle identities are used, tt does not always represent tan⁡θ2\tan\dfrac{\theta}{2}. For instance, in solving the equation sin⁡4θ+tan⁡2θ=0\sin 4\theta + \tan 2\theta = 0 we would use t=tan⁡2θt = \tan 2\theta.

The Expression acos⁡θ+bsin⁡θa\cos\theta + b\sin\theta

It is often useful to reduce acos⁡θ+bsin⁡θa\cos\theta + b\sin\theta to a single term such as rcos⁡(θ−α)r\cos(\theta - \alpha). This is possible provided that we can find values of rr and α\alpha for which

r(cos⁡θcos⁡α+sin⁡θsin⁡α)=acos⁡θ+bsin⁡θr(\cos\theta\cos\alpha + \sin\theta\sin\alpha) = a\cos\theta + b\sin\theta

for every θ\theta. Comparing the coefficients of cos⁡θ\cos\theta and of sin⁡θ\sin\theta,

rcos⁡α=a,rsin⁡α=b.r\cos\alpha = a, \qquad r\sin\alpha = b .

Squaring and adding gives r2=a2+b2r^2 = a^2 + b^2, so rr is equal to the length of the hypotenuse of the triangle containing α\alpha with sides aa and bb, and dividing gives tan⁡α=ba\tan\alpha = \dfrac{b}{a}, with α\alpha in the quadrant given by the signs of aa and bb. Thus

acos⁡θ+bsin⁡θ=rcos⁡(θ−α),wherer=a2+b2andtan⁡α=ba.a\cos\theta + b\sin\theta = r\cos(\theta - \alpha), \qquad\text{where}\qquad r = \sqrt{a^2 + b^2} \quad\text{and}\quad \tan\alpha = \frac{b}{a} .

Example 5.52.

Express 3cos⁡θ+4sin⁡θ3\cos\theta + 4\sin\theta in the form rcos⁡(θ−α)r\cos(\theta - \alpha), giving the values of rr and α\alpha.

Let

r(cos⁡θcos⁡α+sin⁡θsin⁡α)=3cos⁡θ+4sin⁡θ,r(\cos\theta\cos\alpha + \sin\theta\sin\alpha) = 3\cos\theta + 4\sin\theta,

so that rcos⁡α=3r\cos\alpha = 3 and rsin⁡α=4r\sin\alpha = 4. Thus r=5r = 5 and tan⁡α=43\tan\alpha = \dfrac43, which gives α=53.13∘\alpha = 53.13^\circ, and

3cos⁡θ+4sin⁡θ=5cos⁡(θ−53.13∘).3\cos\theta + 4\sin\theta = 5\cos(\theta - 53.13^\circ) .

It is sometimes more convenient to begin by comparing acos⁡θ+bsin⁡θa\cos\theta + b\sin\theta with rsin⁡(θ+α)r\sin(\theta + \alpha), so that

r(sin⁡θcos⁡α+cos⁡θsin⁡α)=acos⁡θ+bsin⁡θ,rsin⁡α=a,rcos⁡α=b.r(\sin\theta\cos\alpha + \cos\theta\sin\alpha) = a\cos\theta + b\sin\theta, \qquad r\sin\alpha = a, \quad r\cos\alpha = b .

Then tan⁡α=ab\tan\alpha = \dfrac{a}{b} and r=a2+b2r = \sqrt{a^2 + b^2}. The value of α\alpha is not the same as it was when we used rcos⁡(θ−α)r\cos(\theta - \alpha). Further variations that could be used are rsin⁡(θ−α)r\sin(\theta - \alpha) and rcos⁡(θ+α)r\cos(\theta + \alpha). When using this method it is better to work from the basic comparison each time, as in the example above, than to quote values of rr and α\alpha.

Problem 5.16.

  1. Express 5cos⁡θ+12sin⁡θ5\cos\theta + 12\sin\theta in the forms rcos⁡(θ−α)r\cos(\theta - \alpha) and rsin⁡(θ+α)r\sin(\theta + \alpha).
  2. Express 3sin⁡θ+4cos⁡θ3\sin\theta + 4\cos\theta in the forms rsin⁡(θ+α)r\sin(\theta + \alpha) and rcos⁡(θ−α)r\cos(\theta - \alpha).
  3. Express 3cos⁡θ−sin⁡θ\sqrt3\cos\theta - \sin\theta in the form rcos⁡(θ+α)r\cos(\theta + \alpha).

The Graph of acos⁡θ+bsin⁡θa\cos\theta + b\sin\theta

First consider the function f(θ)=kcos⁡θf(\theta) = k\cos\theta, where k>0k > 0, which has the following characteristics.

  1. f(θ)=0f(\theta) = 0 when θ=π2,3π2,5π2,…\theta = \dfrac{\pi}{2}, \dfrac{3\pi}{2}, \dfrac{5\pi}{2}, \ldots.
  2. f(θ)=kf(\theta) = k when θ=0,2π,4π,…\theta = 0, 2\pi, 4\pi, \ldots.
  3. f(θ)=−kf(\theta) = -k when θ=π,3π,5π,…\theta = \pi, 3\pi, 5\pi, \ldots.

So the graph of this function is very similar to a standard cosine curve, but has maximum and minimum values ±k\pm k; we say that the curve has an amplitude of kk. We also saw that the graph of cos⁡(θ−α)\cos(\theta - \alpha) is given by moving the graph of cos⁡θ\cos\theta a distance α\alpha to the right. Combining these two modifications of a standard cosine curve, the graph of the function f(θ)=kcos⁡(θ−α)f(\theta) = k\cos(\theta - \alpha) can be sketched.

−5−115θπ/2π3π/22π5 cos θ−3−113θπ/32π/3π4π/35π/32παkk cos(θ − α), drawn with k = 3 and α = π/3
Figure 5.31. A cosine curve with amplitude 55, and a cosine curve with amplitude kk moved a distance α\alpha to the right.

Consider now the function acos⁡θ+bsin⁡θa\cos\theta + b\sin\theta. At first sight its graph is not easy to visualise, but using acos⁡θ+bsin⁡θ=rcos⁡(θ−α)a\cos\theta + b\sin\theta = r\cos(\theta - \alpha) we see that its graph is a cosine curve modified in two ways.

  1. Its maximum and minimum values are ±r\pm r, that is its amplitude is rr.
  2. Its position is a distance α\alpha to the right of the standard curve.

Example 5.53.

Sketch the graph of the function 3cos⁡θ+4sin⁡θ3\cos\theta + 4\sin\theta from −180∘-180^\circ to 180∘180^\circ.

As in the last example, 3cos⁡θ+4sin⁡θ=rcos⁡(θ−α)3\cos\theta + 4\sin\theta = r\cos(\theta - \alpha) with r=5r = 5 and tan⁡α=43\tan\alpha = \dfrac43, so α=53.13∘\alpha = 53.13^\circ. Hence the graph is a cosine curve with an amplitude of 55 and a phase shift of 53.13∘53.13^\circ to the right.

−180°−90°90°180°θ−55(53.13°, 5)
Figure 5.32. The graph of 3cos⁡θ+4sin⁡θ3\cos\theta + 4\sin\theta, with cos⁡θ\cos\theta dashed.

It is interesting to see how the same graph is produced if the alternative form 3cos⁡θ+4sin⁡θ=rsin⁡(θ+α′)3\cos\theta + 4\sin\theta = r\sin(\theta + \alpha') is used. With this approach r=5r = 5 and tan⁡α′=34\tan\alpha' = \dfrac34, so that 3cos⁡θ+4sin⁡θ=5sin⁡(θ+α′)3\cos\theta + 4\sin\theta = 5\sin(\theta + \alpha'), and the graph is a sine curve with an amplitude of 55 and a displacement of α′\alpha' to the left. But since tan⁡α′=cot⁡α\tan\alpha' = \cot\alpha, the angles α′\alpha' and α\alpha are complementary, that is α+α′=π2\alpha + \alpha' = \dfrac{\pi}{2}. We also know that a cosine curve is the same as a sine curve displaced π2\dfrac{\pi}{2} to the left. So a cosine curve moved a distance α\alpha to the right coincides with a sine curve moved a distance α′\alpha' to the left.

So any correct compound angle form of acos⁡θ+bsin⁡θa\cos\theta + b\sin\theta gives a quick method of sketching the graph of that function, and in particular of finding its maximum and minimum values, which for sine and cosine functions are also the greatest and least values.

The Equation acos⁡θ+bsin⁡θ=ca\cos\theta + b\sin\theta = c

One way of solving an equation of this type, using the half angle formulae, has already been used. A compound angle form gives an alternative method. Applied to the equation sin⁡θ+2cos⁡θ=1\sin\theta + 2\cos\theta = 1, solved above with tt, it gives

2cos⁡θ+sin⁡θ=r(cos⁡θcos⁡α+sin⁡θsin⁡α)=rcos⁡(θ−α),2\cos\theta + \sin\theta = r(\cos\theta\cos\alpha + \sin\theta\sin\alpha) = r\cos(\theta - \alpha),

where rcos⁡α=2r\cos\alpha = 2 and rsin⁡α=1r\sin\alpha = 1, that is tan⁡α=12\tan\alpha = \frac12 and r=5r = \sqrt5. Hence

5cos⁡(θ−α)=1,cos⁡(θ−α)=15,θ−α=360n∘±63.43∘,\sqrt5\cos(\theta - \alpha) = 1, \qquad \cos(\theta - \alpha) = \frac{1}{\sqrt5}, \qquad \theta - \alpha = 360n^\circ \pm 63.43^\circ,

from which θ=360n∘±63.43∘+α\theta = 360n^\circ \pm 63.43^\circ + \alpha. But α=arctan⁡12=26.57∘\alpha = \arctan\frac12 = 26.57^\circ, so θ=360n∘+90∘\theta = 360n^\circ + 90^\circ or θ=360n∘−36.87∘\theta = 360n^\circ - 36.87^\circ. Using the values of nn which give θ\theta between 0∘0^\circ and 360∘360^\circ, that is n=0n = 0 and n=1n = 1, we have θ=90∘,323.13∘\theta = 90^\circ, 323.13^\circ.

Example 5.54.

Express 1−sin⁡2θ1+sin⁡2θ\sqrt{\dfrac{1 - \sin 2\theta}{1 + \sin 2\theta}} in terms of tan⁡θ\tan\theta.

Using sin⁡2θ=2t1+t2\sin 2\theta = \dfrac{2t}{1 + t^2}, where t=tan⁡θt = \tan\theta, gives

1−sin⁡2θ=1+t2−2t1+t2=(1−t)21+t2,1+sin⁡2θ=1+t2+2t1+t2=(1+t)21+t2.1 - \sin 2\theta = \frac{1 + t^2 - 2t}{1 + t^2} = \frac{(1 - t)^2}{1 + t^2}, \qquad 1 + \sin 2\theta = \frac{1 + t^2 + 2t}{1 + t^2} = \frac{(1 + t)^2}{1 + t^2} .

Hence

1−sin⁡2θ1+sin⁡2θ=(1−t)2(1+t)2,so1−sin⁡2θ1+sin⁡2θ=±1−tan⁡θ1+tan⁡θ,\frac{1 - \sin 2\theta}{1 + \sin 2\theta} = \frac{(1 - t)^2}{(1 + t)^2}, \qquad\text{so}\qquad \sqrt{\frac{1 - \sin 2\theta}{1 + \sin 2\theta}} = \pm\frac{1 - \tan\theta}{1 + \tan\theta},

taking the sign which makes the right-hand side positive.

Example 5.55.

Find the general solution of the equation cos⁡θ−3sin⁡θ=1\cos\theta - \sqrt3\sin\theta = 1, first by using the half angle formulae and then by using a compound angle form.

  1. With t=tan⁡θ2t = \tan\dfrac{\theta}{2}, the equation becomes 1−t2−23 t1+t2=1\dfrac{1 - t^2 - 2\sqrt3\,t}{1 + t^2} = 1, so 1−t2−23 t=1+t21 - t^2 - 2\sqrt3\,t = 1 + t^2, that is 2t(t+3)=02t(t + \sqrt3) = 0. Hence either t=0t = 0 or t=−3t = -\sqrt3, and the principal values of θ2\dfrac{\theta}{2} are 00 and −π3-\dfrac{\pi}{3}. So θ2=nπ\dfrac{\theta}{2} = n\pi or θ2=nπ−π3\dfrac{\theta}{2} = n\pi - \dfrac{\pi}{3}, and θ=2nπ\theta = 2n\pi or θ=2nπ−2π3\theta = 2n\pi - \dfrac{2\pi}{3}. When tt is undefined, θ=(2n+1)π\theta = (2n + 1)\pi, the left-hand side is −1-1, so no solutions are lost.
  2. Let cos⁡θ−3sin⁡θ=r(cos⁡θcos⁡α−sin⁡θsin⁡α)=rcos⁡(θ+α)\cos\theta - \sqrt3\sin\theta = r(\cos\theta\cos\alpha - \sin\theta\sin\alpha) = r\cos(\theta + \alpha), where rcos⁡α=1r\cos\alpha = 1 and rsin⁡α=3r\sin\alpha = \sqrt3. Then tan⁡α=3\tan\alpha = \sqrt3, so α=π3\alpha = \dfrac{\pi}{3} and r=2r = 2, and the equation can be written 2cos⁡(θ+π3)=12\cos\left(\theta + \dfrac{\pi}{3}\right) = 1. The principal value of θ+π3\theta + \dfrac{\pi}{3} is π3\dfrac{\pi}{3}, so θ+π3=2nπ±π3\theta + \dfrac{\pi}{3} = 2n\pi \pm \dfrac{\pi}{3}, and again θ=2nπ\theta = 2n\pi or θ=2nπ−2π3\theta = 2n\pi - \dfrac{2\pi}{3}.

Example 5.56.

Express 5sin⁡θ+12cos⁡θ5\sin\theta + 12\cos\theta in the form rsin⁡(θ+α)r\sin(\theta + \alpha), giving the values of rr and α\alpha. Show that 5sin⁡θ+12cos⁡θ+7⩽205\sin\theta + 12\cos\theta + 7 \leqslant 20, and find the minimum value of 5sin⁡θ+12cos⁡θ+75\sin\theta + 12\cos\theta + 7. Sketch the graph of the function 15sin⁡θ+12cos⁡θ\dfrac{1}{5\sin\theta + 12\cos\theta} for 0⩽θ⩽2π0 \leqslant \theta \leqslant 2\pi.

Let 5sin⁡θ+12cos⁡θ=r(sin⁡θcos⁡α+cos⁡θsin⁡α)=rsin⁡(θ+α)5\sin\theta + 12\cos\theta = r(\sin\theta\cos\alpha + \cos\theta\sin\alpha) = r\sin(\theta + \alpha), so that rcos⁡α=5r\cos\alpha = 5 and rsin⁡α=12r\sin\alpha = 12, giving r=13r = 13 and tan⁡α=125\tan\alpha = \dfrac{12}{5}, α=67.38∘\alpha = 67.38^\circ. Hence

5sin⁡θ+12cos⁡θ=13sin⁡(θ+α).5\sin\theta + 12\cos\theta = 13\sin(\theta + \alpha) .

But −1⩽sin⁡(θ+α)⩽1-1 \leqslant \sin(\theta + \alpha) \leqslant 1, so −13⩽5sin⁡θ+12cos⁡θ⩽13-13 \leqslant 5\sin\theta + 12\cos\theta \leqslant 13, and adding 77 throughout gives

−6⩽5sin⁡θ+12cos⁡θ+7⩽20.-6 \leqslant 5\sin\theta + 12\cos\theta + 7 \leqslant 20 .

This shows that 5sin⁡θ+12cos⁡θ+7⩽205\sin\theta + 12\cos\theta + 7 \leqslant 20, and that the minimum value of 5sin⁡θ+12cos⁡θ+75\sin\theta + 12\cos\theta + 7 is −6-6.

Now 5sin⁡θ+12cos⁡θ=13sin⁡(θ+α)5\sin\theta + 12\cos\theta = 13\sin(\theta + \alpha), so the graph of 15sin⁡θ+12cos⁡θ\dfrac{1}{5\sin\theta + 12\cos\theta} is also the graph of 113cosec⁡(θ+α)\dfrac{1}{13}\operatorname{cosec}(\theta + \alpha). Its general shape is a typical cosecant curve, except that its branches turn at the values 113\dfrac{1}{13} and −113-\dfrac{1}{13}, and its position is a distance α\alpha to the left of the standard curve.

θπ/2π3π/22π1/13−1/13
Figure 5.33. The graph of 15sin⁡θ+12cos⁡θ\dfrac{1}{5\sin\theta + 12\cos\theta} for 0⩽θ⩽2π0 \leqslant \theta \leqslant 2\pi.

Problem 5.17.

  1. Using t=tan⁡θ2t = \tan\dfrac{\theta}{2}, solve the equation 3cos⁡θ+2sin⁡θ=33\cos\theta + 2\sin\theta = 3, giving the values of θ\theta from −180∘-180^\circ to 180∘180^\circ.
  2. Find the maximum and minimum values of 7cos⁡θ−24sin⁡θ+37\cos\theta - 24\sin\theta + 3, and the values of θ\theta between 0∘0^\circ and 360∘360^\circ at which they occur.
  3. Find the general solution of the equation cos⁡x+sin⁡x=2\cos x + \sin x = \sqrt2.

The Factor Formulae

To factorise is to express in the form of a product. The set of identities called the factor formulae converts expressions such as sin⁡A+sin⁡B\sin A + \sin B into a product. To derive them we use the compound angle group. Adding and subtracting

sin⁡Acos⁡B+cos⁡Asin⁡B=sin⁡(A+B),sin⁡Acos⁡B−cos⁡Asin⁡B=sin⁡(A−B)\sin A\cos B + \cos A\sin B = \sin(A + B), \qquad \sin A\cos B - \cos A\sin B = \sin(A - B)

gives the first two of the following identities, and similar treatment of cos⁡(A+B)\cos(A + B) and cos⁡(A−B)\cos(A - B) gives the other two:

2sin⁡Acos⁡B=sin⁡(A+B)+sin⁡(A−B),(1)2cos⁡Asin⁡B=sin⁡(A+B)−sin⁡(A−B),(2)2cos⁡Acos⁡B=cos⁡(A+B)+cos⁡(A−B),(3)−2sin⁡Asin⁡B=cos⁡(A+B)−cos⁡(A−B).(4)\begin{aligned} 2\sin A\cos B &= \sin(A + B) + \sin(A - B), &\qquad &(1) \\ 2\cos A\sin B &= \sin(A + B) - \sin(A - B), & &(2) \\ 2\cos A\cos B &= \cos(A + B) + \cos(A - B), & &(3) \\ -2\sin A\sin B &= \cos(A + B) - \cos(A - B). & &(4) \end{aligned}

Putting A+B=PA + B = P and A−B=QA - B = Q, so that A=P+Q2A = \dfrac{P + Q}{2} and B=P−Q2B = \dfrac{P - Q}{2}, these become

sin⁡P+sin⁡Q=2sin⁡P+Q2cos⁡P−Q2,(5)sin⁡P−sin⁡Q=2cos⁡P+Q2sin⁡P−Q2,(6)cos⁡P+cos⁡Q=2cos⁡P+Q2cos⁡P−Q2,(7)cos⁡P−cos⁡Q=−2sin⁡P+Q2sin⁡P−Q2.(8)\begin{aligned} \sin P + \sin Q &= 2\sin\frac{P + Q}{2}\cos\frac{P - Q}{2}, &\qquad &(5) \\ \sin P - \sin Q &= 2\cos\frac{P + Q}{2}\sin\frac{P - Q}{2}, & &(6) \\ \cos P + \cos Q &= 2\cos\frac{P + Q}{2}\cos\frac{P - Q}{2}, & &(7) \\ \cos P - \cos Q &= -2\sin\frac{P + Q}{2}\sin\frac{P - Q}{2}. & &(8) \end{aligned}

Identities (5)(5) to (8)(8) are best used when a sum or difference is to be expressed as a product, while identities (1)(1) to (4)(4) should be used when a given product is to be changed into a sum or difference. For example, to express sin⁡6θ−sin⁡4θ\sin 6\theta - \sin 4\theta as a product we use (6)(6):

sin⁡6θ−sin⁡4θ=2cos⁡6θ+4θ2sin⁡6θ−4θ2=2cos⁡5θsin⁡θ.\sin 6\theta - \sin 4\theta = 2\cos\frac{6\theta + 4\theta}{2}\sin\frac{6\theta - 4\theta}{2} = 2\cos 5\theta\sin\theta .

But to express 2cos⁡7θcos⁡2θ2\cos 7\theta\cos 2\theta as a sum we use (3)(3):

2cos⁡7θcos⁡2θ=cos⁡(7θ+2θ)+cos⁡(7θ−2θ)=cos⁡9θ+cos⁡5θ.2\cos 7\theta\cos 2\theta = \cos(7\theta + 2\theta) + \cos(7\theta - 2\theta) = \cos 9\theta + \cos 5\theta .

When these identities are used regularly they are not too difficult to remember. Most people find it best to memorise them in words rather than as symbols. For example, (5)(5) can be remembered as “the sum of two sines is twice the sine of the semi-sum times the cosine of the semi-difference”, and (1)(1) as “twice sin cos is the sine of the sum plus the sine of the difference”. Identities (4)(4) and (8)(8) need special care because of the minus sign.

Example 5.57.

Show that sin⁡A+sin⁡Bcos⁡A+cos⁡B=tan⁡A+B2\dfrac{\sin A + \sin B}{\cos A + \cos B} = \tan\dfrac{A + B}{2}. If AA, BB and CC are the angles of a triangle, deduce that sin⁡A+sin⁡Bcos⁡A+cos⁡B=cot⁡C2\dfrac{\sin A + \sin B}{\cos A + \cos B} = \cot\dfrac{C}{2}.

Considering the left-hand side, by (5)(5) and (7)(7),

sin⁡A+sin⁡Bcos⁡A+cos⁡B=2sin⁡A+B2cos⁡A−B22cos⁡A+B2cos⁡A−B2=tan⁡A+B2.\frac{\sin A + \sin B}{\cos A + \cos B} = \frac{2\sin\frac{A + B}{2}\cos\frac{A - B}{2}}{2\cos\frac{A + B}{2}\cos\frac{A - B}{2}} = \tan\frac{A + B}{2} .

If AA, BB and CC are the angles of a triangle, A+B+C=180∘A + B + C = 180^\circ, so A+B2+C2=90∘\dfrac{A + B}{2} + \dfrac{C}{2} = 90^\circ. The angles A+B2\dfrac{A + B}{2} and C2\dfrac{C}{2} are complementary, therefore tan⁡A+B2=cot⁡C2\tan\dfrac{A + B}{2} = \cot\dfrac{C}{2}, and

sin⁡A+sin⁡Bcos⁡A+cos⁡B=cot⁡C2.\frac{\sin A + \sin B}{\cos A + \cos B} = \cot\frac{C}{2} .

Example 5.58.

Solve the equation sin⁡5x−sin⁡3x=0\sin 5x - \sin 3x = 0, giving the general solution.

By (6)(6),

sin⁡5x−sin⁡3x=2cos⁡5x+3x2sin⁡5x−3x2=2cos⁡4xsin⁡x=0,\sin 5x - \sin 3x = 2\cos\frac{5x + 3x}{2}\sin\frac{5x - 3x}{2} = 2\cos 4x\sin x = 0,

so either cos⁡4x=0\cos 4x = 0 or sin⁡x=0\sin x = 0. If cos⁡4x=0\cos 4x = 0, then 4x=(2n+1)π24x = (2n + 1)\dfrac{\pi}{2}, and if sin⁡x=0\sin x = 0, then x=nπx = n\pi. The general solution is therefore

x=(2n+1)π8orx=nπ.x = (2n + 1)\frac{\pi}{8} \qquad\text{or}\qquad x = n\pi .

Example 5.59.

Factorise cos⁡θ−cos⁡3θ−cos⁡5θ+cos⁡7θ\cos\theta - \cos 3\theta - \cos 5\theta + \cos 7\theta.

Grouping the terms in pairs,

(cos⁡7θ+cos⁡θ)−(cos⁡5θ+cos⁡3θ)=2cos⁡4θcos⁡3θ−2cos⁡4θcos⁡θ=2cos⁡4θ (cos⁡3θ−cos⁡θ),(\cos 7\theta + \cos\theta) - (\cos 5\theta + \cos 3\theta) = 2\cos 4\theta\cos 3\theta - 2\cos 4\theta\cos\theta = 2\cos 4\theta\,(\cos 3\theta - \cos\theta),

and by (8)(8), cos⁡3θ−cos⁡θ=−2sin⁡2θsin⁡θ\cos 3\theta - \cos\theta = -2\sin 2\theta\sin\theta. So

cos⁡θ−cos⁡3θ−cos⁡5θ+cos⁡7θ=−4cos⁡4θsin⁡2θsin⁡θ.\cos\theta - \cos 3\theta - \cos 5\theta + \cos 7\theta = -4\cos 4\theta\sin 2\theta\sin\theta .

Other groupings of the four terms can be used, but arranging them so that pairs of cosines are added makes the factorising simplest.

Example 5.60.

If AA, BB and CC are the angles of a triangle, show that sin⁡A+sin⁡B+sin⁡C=4cos⁡A2cos⁡B2cos⁡C2\sin A + \sin B + \sin C = 4\cos\dfrac{A}{2}\cos\dfrac{B}{2}\cos\dfrac{C}{2}.

Now A+B+C=180∘A + B + C = 180^\circ, so A+B2=90∘−C2\dfrac{A + B}{2} = 90^\circ - \dfrac{C}{2}, and therefore sin⁡A+B2=cos⁡C2\sin\dfrac{A + B}{2} = \cos\dfrac{C}{2} and sin⁡C2=cos⁡A+B2\sin\dfrac{C}{2} = \cos\dfrac{A + B}{2}. Considering the left-hand side, and using (5)(5) and the double angle formula for sin⁡C\sin C,

sin⁡A+sin⁡B+sin⁡C=2sin⁡A+B2cos⁡A−B2+2sin⁡C2cos⁡C2=2cos⁡C2cos⁡A−B2+2cos⁡A+B2cos⁡C2=2cos⁡C2[cos⁡A−B2+cos⁡A+B2].\begin{aligned} \sin A + \sin B + \sin C &= 2\sin\frac{A + B}{2}\cos\frac{A - B}{2} + 2\sin\frac{C}{2}\cos\frac{C}{2} \\ &= 2\cos\frac{C}{2}\cos\frac{A - B}{2} + 2\cos\frac{A + B}{2}\cos\frac{C}{2} \\ &= 2\cos\frac{C}{2}\left[\cos\frac{A - B}{2} + \cos\frac{A + B}{2}\right] . \end{aligned}

But by (7)(7), cos⁡A−B2+cos⁡A+B2=2cos⁡A2cos⁡B2\cos\dfrac{A - B}{2} + \cos\dfrac{A + B}{2} = 2\cos\dfrac{A}{2}\cos\dfrac{B}{2}, so

sin⁡A+sin⁡B+sin⁡C=4cos⁡A2cos⁡B2cos⁡C2.\sin A + \sin B + \sin C = 4\cos\frac{A}{2}\cos\frac{B}{2}\cos\frac{C}{2} .

Problem 5.18.

  1. Factorise sin⁡3A+sin⁡A\sin 3A + \sin A and cos⁡7A−cos⁡A\cos 7A - \cos A.
  2. Solve the equation cos⁡2x+cos⁡4x=0\cos 2x + \cos 4x = 0, giving the values of xx from 0∘0^\circ to 360∘360^\circ.
  3. If AA, BB and CC are the angles of a triangle, show that cos⁡(B+C)=−cos⁡A\cos(B + C) = -\cos A and that cos⁡A+cos⁡B+cos⁡C=1+4sin⁡A2sin⁡B2sin⁡C2\cos A + \cos B + \cos C = 1 + 4\sin\dfrac{A}{2}\sin\dfrac{B}{2}\sin\dfrac{C}{2}.

Small Angles

A glance at the values of sin⁡θ\sin\theta and tan⁡θ\tan\theta when θ\theta is a very small positive angle shows that these two ratios are almost equal. Further, if the small angle is measured in radians, both are found to be almost equal to θ\theta. These relationships can be demonstrated as follows.

Consider a small angle θ\theta, measured in radians, subtended by an arc ABAB at the centre OO of a circle of radius rr. The area of the sector OABOAB is 12r2θ\frac12 r^2\theta. If ACAC is drawn perpendicular to OAOA, to cut OBOB produced at CC, then OACOAC is a right-angled triangle with base rr and height rtan⁡θr\tan\theta, so its area is 12r2tan⁡θ\frac12 r^2\tan\theta. Further, when the chord ABAB is drawn, an isosceles triangle OABOAB is formed with base rr and height rsin⁡θr\sin\theta, so its area is 12r2sin⁡θ\frac12 r^2\sin\theta.

OABCθr
Figure 5.34. The triangle OABOAB, the sector OABOAB and the triangle OACOAC.

Now area of triangle OAB<OAB < area of sector OAB<OAB < area of triangle OACOAC, that is

12r2sin⁡θ<12r2θ<12r2tan⁡θ.\tfrac12 r^2\sin\theta < \tfrac12 r^2\theta < \tfrac12 r^2\tan\theta .

Dividing throughout by 12r2\frac12 r^2, which is positive, gives sin⁡θ<θ<tan⁡θ\sin\theta < \theta < \tan\theta. But sin⁡θ\sin\theta, θ\theta and tan⁡θ\tan\theta are all positive, since θ\theta is a small positive angle, so we can divide throughout by any of them. Dividing by sin⁡θ\sin\theta,

1<θsin⁡θ<sec⁡θ.1 < \frac{\theta}{\sin\theta} < \sec\theta .

For small values of θ\theta, sec⁡θ→1\sec\theta \to 1 as θ→0\theta \to 0. Hence as θ→0\theta \to 0, θsin⁡θ\dfrac{\theta}{\sin\theta} lies between 11 and a number which approaches 11, and we can say that θsin⁡θ→1\dfrac{\theta}{\sin\theta} \to 1 as θ→0\theta \to 0, a limit in the sense of the last lesson. Similarly, by dividing the first inequalities by tan⁡θ\tan\theta, we get cos⁡θ<θtan⁡θ<1\cos\theta < \dfrac{\theta}{\tan\theta} < 1, which shows that θtan⁡θ→1\dfrac{\theta}{\tan\theta} \to 1 as θ→0\theta \to 0.

The same results are obtained if θ\theta is a small negative angle, since θsin⁡θ\dfrac{\theta}{\sin\theta} and θtan⁡θ\dfrac{\theta}{\tan\theta} are unchanged when θ\theta is replaced by −θ-\theta. These limiting values show that, for small values of θ\theta,

sin⁡θ≈θandtan⁡θ≈θ.\sin\theta \approx \theta \qquad\text{and}\qquad \tan\theta \approx \theta .

So far we have not found an approximate value for cos⁡θ\cos\theta when θ\theta is small. To do this we use the double angle identity

cos⁡θ=1−2sin⁡2θ2.\cos\theta = 1 - 2\sin^2\frac{\theta}{2} .

If θ\theta is small then so is θ2\dfrac{\theta}{2}, and sin⁡θ2≈θ2\sin\dfrac{\theta}{2} \approx \dfrac{\theta}{2}, so cos⁡θ≈1−2(θ2)2=1−θ22\cos\theta \approx 1 - 2\left(\dfrac{\theta}{2}\right)^2 = 1 - \dfrac{\theta^2}{2}.

Thus, when θ\theta is measured in radians,

lim⁡θ→0sin⁡θθ=1andlim⁡θ→0tan⁡θθ=1,\lim_{\theta \to 0}\frac{\sin\theta}{\theta} = 1 \qquad\text{and}\qquad \lim_{\theta \to 0}\frac{\tan\theta}{\theta} = 1,

and for any small angle θ\theta, measured in radians,

sin⁡θ≈θ,tan⁡θ≈θ,cos⁡θ≈1−θ22.\sin\theta \approx \theta, \qquad \tan\theta \approx \theta, \qquad \cos\theta \approx 1 - \frac{\theta^2}{2} .

For angles in the range −0.105⩽θ⩽0.105-0.105 \leqslant \theta \leqslant 0.105, that is about −6∘⩽θ⩽6∘-6^\circ \leqslant \theta \leqslant 6^\circ, each of these approximations differs from the true value by less than 0.00050.0005.

Example 5.61.

Find an approximation for the expression sin⁡3θ1+cos⁡2θ\dfrac{\sin 3\theta}{1 + \cos 2\theta} when θ\theta is small.

When θ\theta is small, 3θ3\theta is small, so sin⁡3θ≈3θ\sin 3\theta \approx 3\theta; also 2θ2\theta is small, so cos⁡2θ≈1−(2θ)22=1−2θ2\cos 2\theta \approx 1 - \dfrac{(2\theta)^2}{2} = 1 - 2\theta^2. So when θ\theta is small,

sin⁡3θ1+cos⁡2θ≈3θ2−2θ2.\frac{\sin 3\theta}{1 + \cos 2\theta} \approx \frac{3\theta}{2 - 2\theta^2} .

Example 5.62.

Without using a calculator, find an approximate value for tan⁡61∘\tan 61^\circ, given that 3=1.732\sqrt3 = 1.732 and 1∘=0.0171^\circ = 0.017 radians, giving the answer to three decimal places.

tan⁡61∘=tan⁡(60∘+1∘)=tan⁡60∘+tan⁡1∘1−tan⁡60∘tan⁡1∘=3+tan⁡1∘1−3tan⁡1∘.\tan 61^\circ = \tan(60^\circ + 1^\circ) = \frac{\tan 60^\circ + \tan 1^\circ}{1 - \tan 60^\circ\tan 1^\circ} = \frac{\sqrt3 + \tan 1^\circ}{1 - \sqrt3\tan 1^\circ} .

Now tan⁡θ≈θ\tan\theta \approx \theta when θ\theta is small and measured in radians, so tan⁡1∘≈0.017\tan 1^\circ \approx 0.017. Hence

tan⁡61∘≈1.732+0.0171−(1.732)(0.017)=1.802.\tan 61^\circ \approx \frac{1.732 + 0.017}{1 - (1.732)(0.017)} = 1.802 .

Problem 5.19.

  1. If θ\theta is small, find approximations for θsin⁡θ1−cos⁡θ\dfrac{\theta\sin\theta}{1 - \cos\theta} and sin⁡4θθ\dfrac{\sin 4\theta}{\theta}.
  2. If θ\theta is small enough for θ2\theta^2 to be neglected, show that tan⁡(π4+θ)≈1+θ1−θ\tan\left(\dfrac{\pi}{4} + \theta\right) \approx \dfrac{1 + \theta}{1 - \theta}.

Further Properties of Triangles

There is a great variety of relationships between the sides and angles of a triangle besides the sine and cosine rules, and some of the most useful can be derived from the identities above.

The Cotangent Formula

If DD divides the side ABAB of a triangle ABCABC in the ratio m:nm : n then, with the angles marked in Figure 5.35,

(m+n)cot⁡θ=mcot⁡α−ncot⁡β.(m + n)\cot\theta = m\cot\alpha - n\cot\beta .

This relationship is known as the cotangent formula, and we show it by using the sine rule in the triangles ACDACD and BCDBCD. In the triangle ACDACD the angle at DD is 180∘−θ180^\circ - \theta, so the angle AA is θ−α\theta - \alpha, and

CDsin⁡A=ADsin⁡α,CD=ADsin⁡(θ−α)sin⁡α.\frac{CD}{\sin A} = \frac{AD}{\sin\alpha}, \qquad CD = \frac{AD\sin(\theta - \alpha)}{\sin\alpha} .

In the triangle BCDBCD the angle BB is 180∘−θ−β180^\circ - \theta - \beta, so sin⁡B=sin⁡(θ+β)\sin B = \sin(\theta + \beta), and

CDsin⁡B=BDsin⁡β,CD=BDsin⁡(θ+β)sin⁡β.\frac{CD}{\sin B} = \frac{BD}{\sin\beta}, \qquad CD = \frac{BD\sin(\theta + \beta)}{\sin\beta} .

Hence

ADsin⁡(θ−α)sin⁡α=BDsin⁡(θ+β)sin⁡β.\frac{AD\sin(\theta - \alpha)}{\sin\alpha} = \frac{BD\sin(\theta + \beta)}{\sin\beta} .

But AD:DB=m:nAD : DB = m : n, so

msin⁡(θ−α)sin⁡α=nsin⁡(θ+β)sin⁡β,that ism(sin⁡θcos⁡α−cos⁡θsin⁡α)sin⁡α=n(sin⁡θcos⁡β+cos⁡θsin⁡β)sin⁡β.\frac{m\sin(\theta - \alpha)}{\sin\alpha} = \frac{n\sin(\theta + \beta)}{\sin\beta}, \qquad\text{that is}\qquad \frac{m(\sin\theta\cos\alpha - \cos\theta\sin\alpha)}{\sin\alpha} = \frac{n(\sin\theta\cos\beta + \cos\theta\sin\beta)}{\sin\beta} .

Dividing both sides by sin⁡θ\sin\theta gives

m(cot⁡α−cot⁡θ)=n(cot⁡β+cot⁡θ),so(m+n)cot⁡θ=mcot⁡α−ncot⁡β.m(\cot\alpha - \cot\theta) = n(\cot\beta + \cot\theta), \qquad\text{so}\qquad (m + n)\cot\theta = m\cot\alpha - n\cot\beta .

If DD is the midpoint of ABAB, this becomes 2cot⁡θ=cot⁡α−cot⁡β2\cot\theta = \cot\alpha - \cot\beta.

ABDCθαβmn
Figure 5.35. The angles in the cotangent formula, with AD:DB=m:nAD : DB = m : n.

Other Relationships

The Projection Formula

In any triangle ABCABC,

c=bcos⁡A+acos⁡B.c = b\cos A + a\cos B .

When AA and BB are acute, the perpendicular from CC to ABAB divides ABAB into two parts of lengths bcos⁡Ab\cos A and acos⁡Ba\cos B, as in Figure 5.36. If one of the angles is obtuse, say BB, the foot of the perpendicular lies beyond BB, and the part acos⁡Ba\cos B is negative, so the formula still holds.

The Difference of Two Sides

In any triangle ABCABC,

a−ba+b=tan⁡A−B2tan⁡C2.\frac{a - b}{a + b} = \tan\frac{A - B}{2}\tan\frac{C}{2} .

We show this by using the sine rule in the form asin⁡A=bsin⁡B=k\dfrac{a}{\sin A} = \dfrac{b}{\sin B} = k, say, so that a=ksin⁡Aa = k\sin A and b=ksin⁡Bb = k\sin B. Hence, by the factor formulae,

a−ba+b=k(sin⁡A−sin⁡B)k(sin⁡A+sin⁡B)=2cos⁡A+B2sin⁡A−B22sin⁡A+B2cos⁡A−B2=tan⁡A−B2cot⁡A+B2.\frac{a - b}{a + b} = \frac{k(\sin A - \sin B)}{k(\sin A + \sin B)} = \frac{2\cos\frac{A + B}{2}\sin\frac{A - B}{2}}{2\sin\frac{A + B}{2}\cos\frac{A - B}{2}} = \tan\frac{A - B}{2}\cot\frac{A + B}{2} .

But A+B+C=πA + B + C = \pi, so A+B2=π2−C2\dfrac{A + B}{2} = \dfrac{\pi}{2} - \dfrac{C}{2}, and cot⁡A+B2=tan⁡C2\cot\dfrac{A + B}{2} = \tan\dfrac{C}{2}. So

a−ba+b=tan⁡A−B2tan⁡C2.\frac{a - b}{a + b} = \tan\frac{A - B}{2}\tan\frac{C}{2} .

The Bisector of an Angle

In a triangle ABCABC the bisector of the angle AA divides BCBC in the ratio c:bc : b.

We show this by letting the bisector meet BCBC at DD, with the angle ADBADB equal to θ\theta. In the triangle ABDABD,

BDsin⁡12A=csin⁡θ,\frac{BD}{\sin\frac12 A} = \frac{c}{\sin\theta},

and in the triangle ACDACD, where the angle at DD is 180∘−θ180^\circ - \theta,

DCsin⁡12A=bsin⁡(180∘−θ)=bsin⁡θ.\frac{DC}{\sin\frac12 A} = \frac{b}{\sin(180^\circ - \theta)} = \frac{b}{\sin\theta} .

Hence

sin⁡12Asin⁡θ=BDc=DCb,soBD:DC=c:b.\frac{\sin\frac12 A}{\sin\theta} = \frac{BD}{c} = \frac{DC}{b}, \qquad\text{so}\qquad BD : DC = c : b .
ABCb cos Aa cos Bbac = b cos A + a cos BABCDcbthe bisector of the angle A
Figure 5.36. The perpendicular from CC divides ABAB into bcos⁡Ab\cos A and acos⁡Ba\cos B; the bisector ADAD divides BCBC in the ratio c:bc : b.

Points Associated with a Triangle

The following geometric properties of a triangle should also be familiar.

  1. The perpendicular bisectors of the sides of a triangle ABCABC meet at a point called the circumcentre, which is the centre of the circle through AA, BB and CC, the circumcircle. Its radius is the RR of the sine rule.
  2. The bisectors of the angles AA, BB and CC meet at a point called the incentre, which is the centre of the circle that touches all three sides, the inscribed circle.
  3. The altitudes meet at a point HH called the orthocentre.
  4. The medians meet at a point GG called the centroid.
ABCOcircumcentre OABCIincentre IABCHorthocentre HABCGcentroid G
Figure 5.37. The circumcentre, incentre, orthocentre and centroid of a triangle.

The Area of a Triangle

The area of a triangle can be found using any of the following.

  1. Half the base times the perpendicular height.
  2. 12absin⁡C\frac12 ab\sin C, or the corresponding 12bcsin⁡A\frac12 bc\sin A and 12casin⁡B\frac12 ca\sin B. This follows from the first, since the height of AA above the side CBCB, of length aa, is bsin⁡Cb\sin C.
  3. s(s−a)(s−b)(s−c)\sqrt{s(s - a)(s - b)(s - c)}, where s=12(a+b+c)s = \frac12(a + b + c), which we accept here.

Example 5.63.

In a surveying exercise, PP and QQ are two points on land which is inaccessible. To find the distance PQPQ, a line ABAB of length 200200 metres is drawn so that PP and QQ are on opposite sides of ABAB. The following angles are measured:

∠ABP=60∘,∠ABQ=46∘,∠BAP=30∘,∠BAQ=67∘.\angle ABP = 60^\circ, \qquad \angle ABQ = 46^\circ, \qquad \angle BAP = 30^\circ, \qquad \angle BAQ = 67^\circ .

Find the distance PQPQ.

In the triangle APBAPB, ∠APB=90∘\angle APB = 90^\circ, hence AP=200cos⁡30∘=173.2AP = 200\cos 30^\circ = 173.2 m. In the triangle ABQABQ, ∠AQB=67∘\angle AQB = 67^\circ, hence

AQsin⁡46∘=200sin⁡67∘,AQ=156.3 m.\frac{AQ}{\sin 46^\circ} = \frac{200}{\sin 67^\circ}, \qquad AQ = 156.3 \text{ m}.

Then in the triangle APQAPQ, where ∠PAQ=30∘+67∘=97∘\angle PAQ = 30^\circ + 67^\circ = 97^\circ, the cosine rule gives

PQ2=AP2+AQ2−2⋅AP⋅AQcos⁡97∘=173.22+156.32−2(173.2)(156.3)(−0.1219),PQ^2 = AP^2 + AQ^2 - 2 \cdot AP \cdot AQ\cos 97^\circ = 173.2^2 + 156.3^2 - 2(173.2)(156.3)(-0.1219),

hence PQ=247PQ = 247 m.

ABPQ30°60°67°46°200 m
Figure 5.38. The baseline ABAB and the inaccessible points PP and QQ.

Example 5.64.

In a triangle ABCABC, BC=7.4BC = 7.4 cm, AC=4.1AC = 4.1 cm and ∠ACB=66∘\angle ACB = 66^\circ. Calculate the other angles of the triangle, the area of the triangle, and the distance of the incentre II of the triangle from BCBC.

  1. Using the cosine rule, c2=7.42+4.12−2(7.4)(4.1)cos⁡66∘c^2 = 7.4^2 + 4.1^2 - 2(7.4)(4.1)\cos 66^\circ gives c=6.85c = 6.85 cm. Then the sine rule, 7.4sin⁡A=4.1sin⁡B=6.85sin⁡66∘\dfrac{7.4}{\sin A} = \dfrac{4.1}{\sin B} = \dfrac{6.85}{\sin 66^\circ}, gives A=80.8∘A = 80.8^\circ and B=33.2∘B = 33.2^\circ; AA is acute since a2<b2+c2a^2 < b^2 + c^2.
  2. The area of the triangle is 12absin⁡C=12(7.4)(4.1)sin⁡66∘=13.86 cm2\frac12 ab\sin C = \frac12(7.4)(4.1)\sin 66^\circ = 13.86\ \text{cm}^2.
  3. The incentre II is the centre of the inscribed circle, and is therefore at the same distance rr from all three sides. Joining AIAI, BIBI and CICI divides the triangle into three triangles, AIBAIB, BICBIC and CIACIA, with areas 12cr\frac12 cr, 12ar\frac12 ar and 12br\frac12 br. So the total area of the triangle ABCABC is 12r(a+b+c)=9.175r\frac12 r(a + b + c) = 9.175r. But this area is 13.86 cm213.86\ \text{cm}^2, hence the distance of II from BCBC is r=1.5r = 1.5 cm.
IrABCcab
Figure 5.39. The incentre II at the distance rr from each side, dividing the triangle into three.

Problem 5.20.

  1. Calculate the area of the triangle whose sides are 1111, 1010 and 1515.
  2. In a triangle PQRPQR, p=2.8p = 2.8 m and q=4.5q = 4.5 m. If the area of the triangle is 5.84 m25.84\ \text{m}^2, find the two possible values of rr.
  3. Show that, for any triangle ABCABC, abc=4ΔRabc = 4\Delta R, where Δ\Delta is the area of the triangle and RR is the radius of its circumcircle.
  4. ABCABC is a triangle and DD is the point on BCBC for which BD=DCBD = DC. The angle BADBAD is 20∘20^\circ and the angle CADCAD is 30∘30^\circ. Find the angle ACBACB.

Exercises

Questions marked with an examining board are taken from past A-level papers: JMB is the Joint Matriculation Board, U of L the University of London, C Cambridge, and AEB the Associated Examining Board.

Exercise 5.1.

A chord of a circle subtends an angle of θ\theta radians at the centre of the circle. If the area of the minor segment cut off by the chord is one sixth of the area of the circle, show that

sin⁡θ=θ−π3.\sin\theta = \theta - \frac{\pi}{3} .

Exercise 5.2.

PP and QQ are points on a circle of radius rr, and the chord PQPQ subtends an angle of 2θ2\theta radians at the centre OO. If AA is the area enclosed by the minor arc PQPQ and the chord PQPQ, and BB is the area enclosed by the arc PQPQ and the tangents to the circle at PP and QQ, show that

A−B=r2(2θ−tan⁡θ−sin⁡θcos⁡θ).A - B = r^2(2\theta - \tan\theta - \sin\theta\cos\theta) .

Exercise 5.3.

Three cylinders are placed in contact with each other with their axes parallel. The radii of the cylinders are 33 cm, 44 cm and 55 cm. An elastic band is stretched round the three cylinders so that the plane of the band is perpendicular to the axes of the cylinders. Calculate the length of the part of the band in contact with the largest cylinder. (U of L)

Exercise 5.4.

Without the use of a calculator, find, for each of the following equations, all the solutions in the interval 0∘⩽x⩽180∘0^\circ \leqslant x \leqslant 180^\circ.

  1. cos⁡(x+30∘)=cos⁡(60∘−3x)\cos(x + 30^\circ) = \cos(60^\circ - 3x);
  2. sin⁡(x+20∘)=cos⁡3x\sin(x + 20^\circ) = \cos 3x. (JMB)

Exercise 5.5.

Find the general solution of the equation cos⁡3x+cos⁡x=sin⁡2x\cos 3x + \cos x = \sin 2x. (U of L)

Exercise 5.6.

  1. If sin⁡(θ−α)=ksin⁡(θ+α)\sin(\theta - \alpha) = k\sin(\theta + \alpha), find tan⁡θ\tan\theta in terms of tan⁡α\tan\alpha and kk, and so determine the possible values of θ\theta between 0∘0^\circ and 360∘360^\circ when k=12k = \frac12 and α=150∘\alpha = 150^\circ.
  2. Show, without the use of a calculator, that x=π10x = \dfrac{\pi}{10} satisfies the equation cos⁡3x=sin⁡2x\cos 3x = \sin 2x. By expressing this equation in terms of sin⁡x\sin x and cos⁡x\cos x, show that sin⁡π10\sin\dfrac{\pi}{10} is a root of the equation 4s2+2s−1=04s^2 + 2s - 1 = 0. (C)

Exercise 5.7.

Find all the values of θ\theta in the range 0⩽θ⩽2π0 \leqslant \theta \leqslant 2\pi for which sin⁡θ+sin⁡3θ=cos⁡θ+cos⁡3θ\sin\theta + \sin 3\theta = \cos\theta + \cos 3\theta. (JMB)

Exercise 5.8.

Find, to the nearest minute, the acute angle α\alpha for which 4cos⁡θ−3sin⁡θ=5cos⁡(θ+α)4\cos\theta - 3\sin\theta = 5\cos(\theta + \alpha). Calculate the values of θ\theta in the interval −180∘<θ<180∘-180^\circ < \theta < 180^\circ for which the function f(θ)=4cos⁡θ−3sin⁡θ−4f(\theta) = 4\cos\theta - 3\sin\theta - 4 attains its greatest value, its least value and the value zero. (JMB)

Exercise 5.9.

  1. Show that (sin⁡2θ−sin⁡θ)(1+2cos⁡θ)=sin⁡3θ(\sin 2\theta - \sin\theta)(1 + 2\cos\theta) = \sin 3\theta.
  2. Find the values of xx between 0∘0^\circ and 360∘360^\circ which satisfy the equation sin⁡x+sin⁡2x=sin⁡3x\sin x + \sin 2x = \sin 3x. (C)

Exercise 5.10.

Show that sec⁡x+tan⁡x=tan⁡(π4+x2)\sec x + \tan x = \tan\left(\dfrac{\pi}{4} + \dfrac{x}{2}\right), and deduce a similar expression for sec⁡x−tan⁡x\sec x - \tan x. Hence find in surd form the values of tan⁡π12\tan\dfrac{\pi}{12} and tan⁡5π12\tan\dfrac{5\pi}{12}. (AEB, 1975)

Exercise 5.11.

  1. Show that cos⁡3θ−sin⁡3θ=(cos⁡θ+sin⁡θ)(1−4cos⁡θsin⁡θ)\cos 3\theta - \sin 3\theta = (\cos\theta + \sin\theta)(1 - 4\cos\theta\sin\theta).
  2. Show that if sec⁡A=cos⁡B+sin⁡B\sec A = \cos B + \sin B, then tan⁡2A=sin⁡2B\tan^2 A = \sin 2B and cos⁡2A=tan⁡2(π4−B)\cos 2A = \tan^2\left(\dfrac{\pi}{4} - B\right). (C)

Exercise 5.12.

By expressing sec⁡2x\sec 2x and tan⁡2x\tan 2x in terms of tan⁡x\tan x, or otherwise, solve the equation 2tan⁡x+sec⁡2x=2tan⁡2x2\tan x + \sec 2x = 2\tan 2x, giving all the solutions between −180∘-180^\circ and 180∘180^\circ. (U of L)

Exercise 5.13.

Express 3sin⁡θ−cos⁡θ\sqrt3\sin\theta - \cos\theta in the form Rsin⁡(θ−α)R\sin(\theta - \alpha), where RR is positive. Find all the values of θ\theta in the range 0∘⩽θ⩽360∘0^\circ \leqslant \theta \leqslant 360^\circ which satisfy the equation 4sin⁡θcos⁡θ=3sin⁡θ−cos⁡θ4\sin\theta\cos\theta = \sqrt3\sin\theta - \cos\theta. (JMB)

Exercise 5.14.

  1. Show that (cot⁡θ+cosec⁡θ)2=1+cos⁡θ1−cos⁡θ(\cot\theta + \operatorname{cosec}\theta)^2 = \dfrac{1 + \cos\theta}{1 - \cos\theta}, and hence, or otherwise, solve the equation (cot⁡2θ+cosec⁡2θ)2=sec⁡2θ(\cot 2\theta + \operatorname{cosec} 2\theta)^2 = \sec 2\theta for values of θ\theta between 0∘0^\circ and 180∘180^\circ.
  2. Find the general solution of the equation sin⁡2x+sin⁡3x+sin⁡5x=0\sin 2x + \sin 3x + \sin 5x = 0. (AEB, 1973)

Exercise 5.15.

For θ≢π(mod2π)\theta\not\equiv\pi\pmod{2\pi}, use the formulae expressing sin⁡θ\sin\theta and cos⁡θ\cos\theta in terms of t=tan⁡θ2t = \tan\dfrac{\theta}{2}, or otherwise, to show that

1+sin⁡θ5+4cos⁡θ=(1+t)29+t2.\frac{1 + \sin\theta}{5 + 4\cos\theta} = \frac{(1 + t)^2}{9 + t^2} .

Check the omitted values of θ\theta directly, and deduce that 0⩽1+sin⁡θ5+4cos⁡θ⩽1090 \leqslant \dfrac{1 + \sin\theta}{5 + 4\cos\theta} \leqslant \dfrac{10}{9} for all values of θ\theta. (C)

Exercise 5.16.

  1. Show that cos⁡4x+sin⁡4x=1−12sin⁡22x\cos^4 x + \sin^4 x = 1 - \frac12\sin^2 2x.
  2. Solve, for 0∘⩽x⩽180∘0^\circ \leqslant x \leqslant 180^\circ, the equation sin⁡x+sin⁡5x=sin⁡3x\sin x + \sin 5x = \sin 3x.
  3. Find the general solution of the equation 3cos⁡x+4sin⁡x=23\cos x + 4\sin x = 2. (U of L)

Exercise 5.17.

Show that cosec⁡θ+cot⁡θ=cot⁡θ2\operatorname{cosec}\theta + \cot\theta = \cot\dfrac{\theta}{2}. Hence

  1. deduce the values, in surd form, of cot⁡π8\cot\dfrac{\pi}{8} and cot⁡π12\cot\dfrac{\pi}{12};
  2. express cosec⁡θ+cosec⁡2θ+cosec⁡4θ\operatorname{cosec}\theta + \operatorname{cosec} 2\theta + \operatorname{cosec} 4\theta as the difference of two cotangents;
  3. show, without using a calculator, that cosec⁡4π15+cosec⁡8π15+cosec⁡16π15+cosec⁡32π15=0\operatorname{cosec}\dfrac{4\pi}{15} + \operatorname{cosec}\dfrac{8\pi}{15} + \operatorname{cosec}\dfrac{16\pi}{15} + \operatorname{cosec}\dfrac{32\pi}{15} = 0. (C)

Exercise 5.18.

  1. Express 7sin⁡x−24cos⁡x7\sin x - 24\cos x in the form Rsin⁡(x−α)R\sin(x - \alpha), where RR is positive and α\alpha is an acute angle. Hence, or otherwise, solve the equation 7sin⁡x−24cos⁡x=157\sin x - 24\cos x = 15 for 0∘⩽x⩽360∘0^\circ \leqslant x \leqslant 360^\circ.
  2. Solve the simultaneous equations cos⁡x+cos⁡y=1\cos x + \cos y = 1 and sec⁡x+sec⁡y=4\sec x + \sec y = 4 for 0∘⩽x⩽180∘0^\circ \leqslant x \leqslant 180^\circ and 0∘⩽y⩽180∘0^\circ \leqslant y \leqslant 180^\circ. (AEB, 1976)

Exercise 5.19.

Express cos⁡2x−sin⁡2x\cos 2x - \sin 2x in the form Rcos⁡(2x+α)R\cos(2x + \alpha), giving the values of RR and α\alpha. Hence find the general solution of each of the following equations.

  1. cos⁡2x−sin⁡2x=1\cos 2x - \sin 2x = 1;
  2. cos⁡2x−sin⁡2x=2cos⁡4x\cos 2x - \sin 2x = \sqrt2\cos 4x. (U of L)

Exercise 5.20.

  1. Find, in the range −180∘⩽x⩽180∘-180^\circ \leqslant x \leqslant 180^\circ, the solutions of the equation cos⁡5x=cos⁡x\cos 5x = \cos x.
  2. Show that 1+cos⁡θ+sin⁡θ1−cos⁡θ+sin⁡θ=1+cos⁡θsin⁡θ\dfrac{1 + \cos\theta + \sin\theta}{1 - \cos\theta + \sin\theta} = \dfrac{1 + \cos\theta}{\sin\theta}. (JMB)

Exercise 5.21.

If sin⁡θ+sin⁡2θ+sin⁡3θ+sin⁡4θ=0\sin\theta + \sin 2\theta + \sin 3\theta + \sin 4\theta = 0, show that θ\theta is either a multiple of π2\dfrac{\pi}{2} or a multiple of 2π5\dfrac{2\pi}{5}. (U of L)

Exercise 5.22.

Write down the expansions of cos⁡(A+B)\cos(A + B) and cos⁡(A−B)\cos(A - B) in terms of the cosines and sines of AA and BB.

  1. Find angles xx and yy, each between 00 and 90∘90^\circ, which satisfy the simultaneous equations cos⁡xcos⁡y=0.6\cos x\cos y = 0.6 and sin⁡xsin⁡y=0.2\sin x\sin y = 0.2.
  2. Show that cos⁡3x=4cos⁡3x−3cos⁡x\cos 3x = 4\cos^3 x - 3\cos x. Hence find all the solutions, in the range −180∘⩽x⩽180∘-180^\circ \leqslant x \leqslant 180^\circ, of the equation 2cos⁡3x+cos⁡2x+1=02\cos 3x + \cos 2x + 1 = 0. (JMB)

Exercise 5.23.

By means of the substitution tan⁡θ=t\tan\theta = t, or otherwise, find the values of θ\theta in the range 0<θ<π20 < \theta < \dfrac{\pi}{2} such that

(2−tan⁡θ)(1+sin⁡2θ)−2=0.(2 - \tan\theta)(1 + \sin 2\theta) - 2 = 0 .

Show that, when θ\theta is small, (2−tan⁡θ)(1+sin⁡2θ)−2≈3θ(2 - \tan\theta)(1 + \sin 2\theta) - 2 \approx 3\theta. (JMB)

Exercise 5.24.

A quadrilateral ABCDABCD is right-angled at BB and at DD, and the angle DABDAB is 132∘132^\circ. If DA=4DA = 4 and AB=7AB = 7, find the lengths of the diagonals of the quadrilateral and the radius of the inscribed circle of the triangle ABCABC. (U of L)

Exercise 5.25.

Three towns AA, BB and CC are all at sea level. The bearings of the towns BB and CC from AA, measured clockwise from north, are 36∘36^\circ and 247∘247^\circ respectively. If BB is 120120 km from AA and CC is 234234 km from AA, calculate the distance and the bearing of the town BB from CC. (AEB, 1972)

Check Yourself

 

Fresh questions on the whole chapter — none of them is worked out above. Do each on paper first; the box only tells you whether you got there.

Answers are checked in your browser, as often as you like. Nothing is sent anywhere and nothing is kept but your own work. A formula may be written with the symbols themselves or with ~ & | -> <-> ^, and \and, \or, \to expand as you type.

Exercise 5.26.

What is 150∘150^\circ in radians?

answer one of these

Exercise 5.27.

An arc of a circle of radius 66 cm subtends an angle of π3\dfrac{\pi}{3} at the centre. What is the length of the arc, in centimetres?

answer one of these

Exercise 5.28.

A sector of a circle of radius 44 cm contains an angle of 1.51.5 radians. What is its area, in square centimetres?

answer one of these

Exercise 5.29.

What is sin⁡210∘\sin 210^\circ?

answer one of these

Exercise 5.30.

The angle θ\theta is obtuse and sin⁡θ=45\sin\theta = \dfrac45. What is cos⁡θ\cos\theta?

answer one of these

Exercise 5.31.

A triangle has sides 55, 77 and 88. What is the angle opposite the side of length 77?

answer one of these

Exercise 5.32.

In a triangle ABCABC, A=30∘A = 30^\circ, B=45∘B = 45^\circ and a=5a = 5. What is bb?

answer one of these

Exercise 5.33.

What is the period of the function θ↦tan⁡3θ\theta \mapsto \tan 3\theta?

answer one of these

Exercise 5.34.

What is arccos⁡(−12)\arccos\left(-\dfrac12\right)?

answer one of these

Exercise 5.35.

What is the general solution of tan⁡θ=13\tan\theta = \dfrac{1}{\sqrt3}?

answer one of these

Exercise 5.36.

How many solutions has the equation sin⁡2θ=12\sin 2\theta = \dfrac12 in the interval 0⩽θ⩽2π0 \leqslant \theta \leqslant 2\pi?

answer one of these

Exercise 5.37.

What is sin⁡15∘cos⁡15∘\sin 15^\circ\cos 15^\circ?

answer one of these

Exercise 5.38.

What is tan⁡105∘\tan 105^\circ?

answer one of these

Exercise 5.39.

What is the greatest value of 3sin⁡θ+4cos⁡θ+13\sin\theta + 4\cos\theta + 1?

answer one of these

Exercise 5.40.

What is the area of the triangle with a=6a = 6, b=8b = 8 and C=30∘C = 30^\circ?

answer one of these

Exercise 5.41.

What is lim⁡θ→0sin⁡3θθ\displaystyle\lim_{\theta \to 0}\frac{\sin 3\theta}{\theta}, with θ\theta in radians?

answer one of these