© 202
5
LSE All Rights Reserved
MODULE 6 UNIT 1 Regression analysis
© 202
5
LSE All Rights Reserved getsmarter.com
|
info@getsmarter.com +44 203 457 5774 (UK)
|
+1 224 249 3522 (US)
|
+27 21 447 7565 (SA) Page 4 of 25
lse.ac.uk
Table 1:
Number of people unemployed and number of low-cost houses in various parts of a city.
Number of unemployed
2,655 1,500 1,489 2,148 3,158 1,780 1,211 1,958 2,478
Number of low-cost houses
920 452 395 756 1,179 524 300 720 898 The first step in determining the relationship between these two variables is to create a
scatterplot
(sometimes called a scatter diagram). This allows you to visually inspect whether the variables are related. If you plot the values provided in Table 1, with the number of people unemployed on the
x
-axis and the number of low-cost houses on the
y
-axis, you get the scatterplot shown in Figure 1.
Figure 1:
A scatterplot of the number of low-cost houses and people unemployed.
Figure 1 gives the impression that there is some sort of linear relationship between the number of people unemployed and the number of low-cost houses in the various parts of the city. The scatterplot shows that, as the number of people unemployed increases, the number of low-cost houses also increases, almost forming a straight line. However, the relationship is not perfectly linear, as the points do not lie on a straight line, and a certain amount of scatter exists.
© 202
5
LSE All Rights Reserved getsmarter.com
|
info@getsmarter.com +44 203 457 5774 (UK)
|
+1 224 249 3522 (US)
|
+27 21 447 7565 (SA) Page 5 of 25
lse.ac.uk
2.1 Types of correlation
When one variable is expected to increase as the other increases, as in Figure 1, it is called
positive correlation
. Inversely, if one variable is shown to decrease as the other increases, it is known as
negative correlation
. When two variables show no signs of linear correlation, they are
uncorrelated
. This means that there is no linear relationship between the two variables. However, it does not necessarily mean that there is no relationship; the variables might be exponentially or quadratically related, for example. These types of correlation will not be covered in this course – only linear correlation will be addressed. Figures 2 to 4 presents some examples of correlation. Figure 2 is an example of positively-correlated data, where an increase in marketing expenditure has resulted in an increase in sales (note that while you are only considering correlation at the moment, you could plausibly infer causality, i.e. marketing expenditure affects total sales, rather than vice versa).
Figure 2:
Total sales in relation to marketing expenditure – positive correlation.
© 202
5
LSE All Rights Reserved getsmarter.com
|
info@getsmarter.com +44 203 457 5774 (UK)
|
+1 224 249 3522 (US)
|
+27 21 447 7565 (SA) Page 6 of 25
lse.ac.uk
Figure 3 shows the negative correlation between the total sales of a product and its price.
Figure 3:
Total sales in relation to price – negative correlation.
Figure 4 shows two variables, namely customers’ shoe sizes and their monthly income, which have no evident correlation. Thus, you could conclude that they are uncorrelated. However, it is possible that they are correlated in some other way, just not linearly; although, based on Figure 4, there is no obvious pattern in the scatter.
© 202
5
LSE All Rights Reserved getsmarter.com
|
info@getsmarter.com +44 203 457 5774 (UK)
|
+1 224 249 3522 (US)
|
+27 21 447 7565 (SA) Page 7 of 25
lse.ac.uk
Figure 4:
Income per month in relation to shoe size – no correlation.
Correlation is measured as a value between -1 and 1. This value is called the
correlation coefficient
. A value of -1 indicates a strong,
negative
, linear relationship between the
x
and
y
variables. Similarly, a value of 1 indicates a strong,
positive
, linear relationship between
x
and
y
. A correlation coefficient of 0 indicates that
x
and
y
are
uncorrelated
, meaning they are not linearly related. The calculation of the correlation coefficient by hand is explored in Section 3 of these notes. The correlation coefficient calculation in Microsoft Excel, and the interpretation thereof, will be discussed in Unit 3 of this module. You will often hear statements made about the correlation between variables, but how can you be certain of such stated relationships? Consider the following variables and the correlation they are expected to have.
Table 2:
Variables and their expected correlation.
Variables Expected correlation
The height and weight of a person Positive correlation Number of hours of sunshine and rainfall Negative correlation Ice cream sales and sunscreen sales Positive correlation Time spent studying and test grades Positive correlation Education and income Positive correlation Absence from work and productivity Negative correlation Petrol consumption and ice cream sales No correlation
© 202
5
LSE All Rights Reserved getsmarter.com
|
info@getsmarter.com +44 203 457 5774 (UK)
|
+1 224 249 3522 (US)
|
+27 21 447 7565 (SA) Page 8 of 25
lse.ac.uk
It is important to note that a correlation existing between variables does not necessarily mean that they are dependent on one another, or that one’s values are the cause of the other's.
2.2 Causality
When two variables have a linear relationship, you may wonder if they also have a
causal
relationship. Suppose the following three examples of variables have been shown to have strong correlations: 1.Number of calories eaten and weight gained.2.A country’s GDP and average life expectancy of its citizens.3.Average salary of teachers and alcohol consumption by country.Would it be reasonable to think that there is a
causal
relationship between these variables? 1.It would be perfectly reasonable to conclude that a person’s weight is influencedby their diet and, as a result, that a person is likely to gain more weight on a high-calorie diet. Therefore, you could conclude that a causal correlation exists betweennumber of calories eaten and weight gained.2.It is tricky to establish if GDP is responsible for life expectancy, so it is difficult toestablish a causal relationship. However, it could be argued that richer countrieshave a population that lives longer because their citizens have access to moreadvanced healthcare and other basic services, meaning that there could be acausal relationship between a country’s GDP and the average life expectancy of its citizens.3.It would be frivolous to conclude that increasing the remuneration of teacherswould increase alcohol consumption. It is much more plausible that both variablesare the result of various factors – including the country’s economy and its socialvalues – and probably do not have a causal relationship.It should be clear, at this point, that you should take care when interpreting correlated relationships, as
correlation does not imply causality
.
Explore further:
Tyler Vigen humorously emphasises that correlation does not equal causality on his website Spurious Correlations, which shows correlations between unrelated variables.
© 202
5
LSE All Rights Reserved getsmarter.com
|
info@getsmarter.com +44 203 457 5774 (UK)
|
+1 224 249 3522 (US)
|
+27 21 447 7565 (SA) Page 9 of 25
lse.ac.uk
3.Regression analysis
Once a linear relationship has been identified between two variables by assessing the correlation coefficient, a regression analysis can be performed to explain this relationship. There are three main reasons a regression analysis would be performed:
●
To establish and interpret the unknown variables in a known relationship
●
To understand the reasons for such a relationship and determine causality
●
To forecast the response of one variable because of another
3.1 Variables
In a simple linear regression model, two variables are present: the
dependent (or response) variable,
and the
independent (or explanatory) variable
.
●
y
is the dependent (or response) variable, meaning the factor that you are trying toexplain.
●
x
is the independent (or explanatory) variable, meaning the factor that is thoughtto have an influence on
y
.In the earlier example of the number of people unemployed and the corresponding number of low-cost houses, it could be argued that unemployment is what drives the need for low-cost housing. Using this argument, unemployment levels would be the explanatory variable which explains the response variable, low-cost housing.
3.2 Summary statistics
In order to draw inferences from the outcome of regression analyses, some summary statistics need to be defined first. For a paired dataset – where the linear relationship has been determined and the dependent and independent variables have been defined – the corrected sum of squares and products need to be calculated.
Note:
These summary statistics are automatically calculated when performing a regression analysis using built-in functions, such as those found in Microsoft Excel. However, in order to understand how they are calculated, and to interpret them, their equations are discussed in this section. The sum of squares is calculated using the squared distance that each data point is situated away from the mean of that variable. In other words, how far away any
x
value is from
x̄
, or how far away any
y
value is from
ȳ
.
© 202
5
LSE All Rights Reserved getsmarter.com
|
info@getsmarter.com +44 203 457 5774 (UK)
|
+1 224 249 3522 (US)
|
+27 21 447 7565 (SA) Page 10 of 25
lse.ac.uk
The sum of the squares of these values is known as the
corrected sum of squares for
x
-data (
S
xx
)
and the
corrected sum of squares for
y
-data (
S
yy
)
. Another summary statistic is the
corrected sum of cross-products of
x
-data and
y
-data (
S
xy
).
This is calculated by summing the products of the distance the
x
-value is from its mean with the distance the
y
-value is from its mean.
Example 1:
Consider again the following paired dataset, where the number of people unemployed was positively correlated to the number of low-cost houses in various parts of a city. The number of low-cost houses has been defined as the response variable, as it could be argued that it is dependent on unemployment.
Table 3:
Unemployment and reported low-cost housing data.
Number of unemployed (
x
)
2,655 1,500 1,489 2,148 3,158 1,780 1,211 1,958 2,478
Number of low-cost houses (
y
)
920 452 395 756 1,179 524 300 720 898 The means of the
x-
data and
y
-data can be calculated as
x̄
= 2,041.89 and
ȳ
= 682.67, respectively. The corrected sums of squares and corrected sum of cross-products can then be calculated.
Table 4:
Corrected sum of squares and cross-products.
x y
(
x
-
x̄
)
2
(
y
-
ȳ
)
2
(
x
-
x̄
)(
y
-
ȳ
)
2655 920 375,905.23 56,327.11 145,511.70 1,500 452 293,643.57 53,207.11 124,995.70 1,489 395 305,686.12 82,752.11 159,047.70
© 202
5
LSE All Rights Reserved getsmarter.com
|
info@getsmarter.com +44 203 457 5774 (UK)
|
+1 224 249 3522 (US)
|
+27 21 447 7565 (SA) Page 11 of 25
lse.ac.uk
2,148 756 11,259.57 5,377.78 7,781.48 3,158 1,179 1,245,704.01 246,346.78 553,963.15 1,780 524 68,585.79 25,175.11 41,553.04 1,211 300 690,376.35 146,433.78 317,953.48 1,958 720 7,037.35 1,393.78 -3,131.852,478 898 190,192.90 46,368.44 93,909.26
Σ
3,188,390.89 663,382.00 1,441,583.67
These calculated values will now be used in further calculations to determine the correlation and regression components.
3.3 Correlation coefficient
Having calculated the summary statistics of
S
xx
,
S
yy
, and
S
xy,
these can be used to determine the correlation coefficient,
r
. Recall from Section 2.1 that the correlation coefficient is a value between -1 and 1, indicating the strength of the linear relationship between two measurable variables. The equation for the correlation coefficient is:
Example 2:
Looking once more at the unemployment and low-cost housing data in the previous example, where
S
xx
= 3,188,390.89,
S
yy
= 663,382.00, and
S
xy
= 1,441,583.67, the correlation coefficient can be calculated: With
r
being very close to positive 1 (at 0.991), it can be said that there is a strong, positive, linear relationship between the number of people unemployed and the number of low-cost houses in various parts of the city.
3.4 Regression line
Let’s summarise the different variables that have been defined up to this point.
●
x
= the independent (or explanatory) variable. This variable is not dependent onanother and can be changed or chosen as needed. The mean of
x
is
x̄
.
●
y
= the dependent (or response) variable. This variable is expected to change asthe explanatory variable changes. The mean of
y
is
ȳ
.
© 202
5
LSE All Rights Reserved getsmarter.com
|
info@getsmarter.com +44 203 457 5774 (UK)
|
+1 224 249 3522 (US)
|
+27 21 447 7565 (SA) Page 12 of 25
lse.ac.uk
●
S
xx
= corrected sum of squares for
x
-data. This value is calculated by summing thesquares of the distances of all the various values of
x
in relation to
x̄
.
●
S
yy
= corrected sum of squares for
y
-data. This value is calculated by summing thesquares of the distances of all the various values of
y
in relation to
ȳ
.
●
S
xy
= corrected sum of cross-products for
x
-data and
y
-data.
●
r
= the correlation coefficient. A value between -1 and 1, which indicates whether there is a positive or negative linear relationship between the
x
and
y
variables, aswell as the strength of the relationship.For two variables, when a correlation has been identified, the nature of their relationship can be represented using an estimated regression line. As with any straight line, the equation for the estimated regression line uses the standard linear equation of a
y
-intercept summed with the product of the gradient and the
x
-value. In this case, the
y
-intercept is defined as
β
0
and the gradient as
β
1
. This produces the equation: Where
β
0
and
β
1
are unknown parameters that are estimated using the following equations: As this is an approximation of the true relationship between the explanatory and response variables, each
y
observation requires some random digression from the initial approximated line. Each observation of
y
almost lies on the approximated line, with a certain offset distance. This perturbation is called the
error term
and is denoted with the Greek letter epsilon,
ε
. Therefore, the basic model is modified to include the error term, and becomes: The error terms corresponding to the data points are assumed to be independent and identically normally distributed, with a mean of 0 and a constant, unknown variance. This means that any of the individual error terms has a high probability of being small and relatively close to the regression line.
© 202
5
LSE All Rights Reserved getsmarter.com
|
info@getsmarter.com +44 203 457 5774 (UK)
|
+1 224 249 3522 (US)
|
+27 21 447 7565 (SA) Page 13 of 25
lse.ac.uk
Figure 5:
Normally distributed error terms.
Example 3:
Consider once more the data regarding the number of people unemployed and the related number of low-cost houses, and recall the calculated summary statistics of
S
xx
= 3,188,390.89,
S
yy
= 663,382.00, and
S
xy
= 1,441,583.67. Earlier on, the means were calculated as
x̄
= 2,041.89 and
ȳ
= 682.67. Substituting these values into the beta equations, the equation for the estimated regression line can be calculated: Substituting the beta-hat values, the estimated regression line equation becomes: When the
x
and
y
variables are plotted (view Figure 1 for a reminder) and a trendline is included in Excel, the graph shown in Figure 6 is generated.
© 202
5
LSE All Rights Reserved getsmarter.com
|
info@getsmarter.com +44 203 457 5774 (UK)
|
+1 224 249 3522 (US)
|
+27 21 447 7565 (SA) Page 14 of 25
lse.ac.uk
Figure 6:
Regression line created in Excel for low-cost houses and unemployment.
The regression line is useful in that it contains the
beta coefficient
for a paired dataset, also known as the gradient of the line. It only indicates the best-fitting straight line, and the variance in the dataset should be determined and analysed next. Note that the calculated regression line paves the way for predictions made from regression models. These methods of prediction are discussed in Module 7.
3.5 Analysis of variance (ANOVA)
The overall objective of using regression analysis is to explain the response variable as a function of the explanatory variable. For example, for every unit increase spent on marketing, what will the increase in sales figures be? More specifically, the goal is to explain the
variation
in the response variable. When performing a simple linear regression, this is done by using a single explanatory variable. The total variation in the response variable for a paired dataset is the corrected sum of squares of the
y
-data: This can now be called the
total sum of squares
(
TSS
). The total sum of squares can be split into two parts: the first is the part of the
TSS
that can be explained using the model,
© 202
5
LSE All Rights Reserved getsmarter.com
|
info@getsmarter.com +44 203 457 5774 (UK)
|
+1 224 249 3522 (US)
|
+27 21 447 7565 (SA) Page 15 of 25
lse.ac.uk
called the
explained sum of squares
(
ESS
); and the second, which cannot be explained, called the
residual sum of squares
(
RSS
).
TSS = ESS + RSS
Figure 7 shows how the total variation consists of the explained variation and the residual variation.
Figure 7:
Explained and residual variations as parts of total variation.
The
TSS
will thus be the explained sum of squares combined with the residual sum of squares to form the total sum of squares. The equation for calculating the
ESS
is: The overall fit of the regression model can be expressed as the proportion of the total variability in the response variable that is explained by the model. In other words, how much of the total sum of squares is contained in the explained sum of squares? This is known as the
coefficient of determination,
R
2
.
© 202
5
LSE All Rights Reserved getsmarter.com
|
info@getsmarter.com +44 203 457 5774 (UK)
|
+1 224 249 3522 (US)
|
+27 21 447 7565 (SA) Page 16 of 25
lse.ac.uk
As
R
2
is a proportion, it is a value between 0 and 1, and the closer
R
2
is to 1 the better the
explanatory power of the model
.
Note:
Although residuals and error terms are closely related and easily confused, take note of the following fundamental difference between them.
●
Error term: A "correction" term to reflect the fact that the true relationship between
x
and
y
is not exactly linear. It measures the difference between the observed valueand the true value.
●
Residual: The difference between an observed value of
y
and its fitted value fromthe estimated regression line.
Example 4:
Continuing once more with the example of the number of people unemployed and the related number of low-cost houses, recall the calculated summary statistics of
S
xx
= 3,188,390.89,
S
yy
= 663,382.00, and
S
xy
= 1,441,583.67. Calculating
R
2
: This means that 98.25% of the total variation in the number of low-cost houses can be explained by unemployment. Note that this is extremely high, indicating that the model is a great fit.
3.6 Multivariate analysis
Multivariate analysis is used when a response variable is expected to have a linear relationship with multiple explanatory variables. These relationships are not usually equal, as one explanatory variable can have a greater influence on the response variable than another. For two explanatory variables, the resultant linear regression line equation obtains another factor, becoming: As more explanatory variables are included, more factors are added to the equation.
© 202
5
LSE All Rights Reserved getsmarter.com
|
info@getsmarter.com +44 203 457 5774 (UK)
|
+1 224 249 3522 (US)
|
+27 21 447 7565 (SA) Page 17 of 25
lse.ac.uk
4.The challenges of linearity
Not all paired datasets are created equal. Remember that for regression to achieve valid and reliable results, the relationship between the explanatory and response variables needs to be at least approximately linear. This is not always the case for relationships between variables. Take, for example, the two variables plotted on a scatter diagram in Figure 8. In this case, the relationship between them is represented with the equation
y = x(1−x)
for values of
x
between 0 and 1.
Figure 8:
Non-linear variables.
Clearly this relationship is not linear. In fact, it follows a parabolic curve. Although not linear, a well-defined, non-linear relationship still clearly exists between the two variables, so they are certainly not independent. For interest's sake, the correlation coefficient for this pair is basically 0, at 5.37 × 10
-17
, resulting in an uncorrelated outcome. It is often useful to explore other ways of representing the same data, only in a more linear form. By transforming an exponential dataset into its logarithmic equivalent, it is transformed into a more linear dataset that does fit the assumptions needed for regression analysis. Regression analysis relies on assumptions made about the linearity of a paired dataset. In Unit 2, you will delve deeper into what these assumptions are and how you can ensure that your regression analyses present valid and reliable results by investigating this linearity.
© 202
5
LSE All Rights Reserved getsmarter.com
|
info@getsmarter.com +44 203 457 5774 (UK)
|
+1 224 249 3522 (US)
|
+27 21 447 7565 (SA) Page 18 of 25
lse.ac.uk
5.Examples
Consider two examples of regression analysis, the first with two variables (bivariate analysis), and the second with multiple variables (multivariate analysis.)
5.1 Bivariate analysis
Working as a marketing manager, Simone is asked to determine whether the amount the company spent on marketing in the last year was related to the total number of sales generated by the company. She is also asked to investigate the nature of the relationship, should one exist. Simone obtains the marketing expenditure and total sales figures for the last year, as shown in Table 5.
Table 5:
Marketing expenditure and total sales for 12 months.
Month Marketing expenditure Total sales
January £1,550 £11,874 February £980 £8,123 March £1,225 £8,578 April £845 £7,015 May £1,125 £8,746 June £1,024 £7,210 July £784 £6,758 August £1,358 £10,045 September £1,204 £10,455 October £1,050 £8,987 November £1,402 £11,874 December £905 £6,251 Simone determines from this paired dataset that total sales could be dependent on the marketing expenditure. Therefore, she defines the marketing expenditure as the independent variable,
x
, and total sales as the dependent variable,
y
. A scatterplot of this data indicates that there seems to be a positive linear relationship between these variables, as shown in Figure 9.
© 202
5
LSE All Rights Reserved getsmarter.com
|
info@getsmarter.com +44 203 457 5774 (UK)
|
+1 224 249 3522 (US)
|
+27 21 447 7565 (SA) Page 19 of 25
lse.ac.uk
Figure 9:
Total sales as dependent variables of marketing expenditure.
Simone calculates the summary statistics as:
x̄
= £1,121
ȳ
= £8,826
S
xx
= 607,624.00
S
yy
= 40,106,668.67
S
xy
= 4,559,474.00 Using these values, she calculates the correlation coefficient as 0.9236, indicating that there is indeed a strong, positive, linear relationship between marketing expenditure and total sales. The regression line components are calculated as: Substituting the beta coefficient estimates, the estimated regression line equation becomes:
© 202
5
LSE All Rights Reserved getsmarter.com
|
info@getsmarter.com +44 203 457 5774 (UK)
|
+1 224 249 3522 (US)
|
+27 21 447 7565 (SA) Page 20 of 25
lse.ac.uk
Plotting the regression line on the scatterplot, Simone generates the graph seen in Figure 10.
Figure 10:
Regression line of total sales in relation to marketing expenditure.
Lastly, Simone calculates the coefficient of determination,
R
2
: Simone reports back to her boss, indicating that the total sales generated by the company can be attributed, in some way, to marketing expenditure. She indicates that 85.31% of the variation in total sales can be explained by marketing expenditure. The high value indicates that the model is a very good fit. She also notes that not all the sales can be explained by marketing expenditure, as is expected, as other factors will also influence how well the company performs.
5.2 Multivariate analysis
Harriet works for a fast-food company and is asked to illustrate the effects of marketing instruments on the weekly sales volumes of a certain product over a three-year period. The marketing instruments used by Harriet’s company were:
●
Changing the promotion price
●
Varying the amount spent on feature advertising
© 202
5
LSE All Rights Reserved getsmarter.com
|
info@getsmarter.com +44 203 457 5774 (UK)
|
+1 224 249 3522 (US)
|
+27 21 447 7565 (SA) Page 21 of 25
lse.ac.uk
●
Varying the amount spent on display measuresHarriet decides to plot the sales volumes over time, and generates the chart shown in Figure 11.
Figure 11:
Total sales volumes per week.
The time series plot shows that there is momentum in the data. It also shows that there are major fluctuations in total sales volumes, which might be due to any of the three marketing instruments. Because of the large fluctuation in sales volumes, Harriet decides to plot the logarithmic total sales volumes against the same time series, which generates the chart provided in Figure 12.
© 202
5
LSE All Rights Reserved getsmarter.com
|
info@getsmarter.com +44 203 457 5774 (UK)
|
+1 224 249 3522 (US)
|
+27 21 447 7565 (SA) Page 22 of 25
lse.ac.uk
Figure 12:
Logarithmic total sales volumes per week.
Immediately, she notices smaller fluctuations in sales volumes because of a more linear time series. Next, Harriet plots the logarithmic total sales volumes against each of the three marketing instruments. From the chart in Figure 13, Harriet notices that there is a negative correlation between promotion price and logarithmic total sales volumes. This makes logical sense, as an increase in price would normally result in decreased sales.