OLS Fundamentals
This page explains the mathematics behind the results shown in the Linear Regression tab. Its central subject is the assumptions OLS places on the errors and which properties of the results each assumption guarantees. See the Linear Regression tab for the tab's operation and how to read the results screen.
Model Formulation and the OLS Estimator
The linear regression model expresses the response as the sum of a linear combination of the predictors and an error:
is the vector of the response values for observations, and is the design matrix. is the coefficient vector and is the error vector. In models with an intercept, the first column of is a constant column of ones, and counts the columns including the intercept. The error is an unobservable quantity: the part of the response's variation that the linear combination of the predictors cannot express.
OLS (ordinary least squares) takes as the coefficient estimate the that minimizes the residual sum of squares . Because the residual sum of squares is quadratic in , the minimizer is unique whenever the columns of are linearly independent, and setting the derivative to zero gives it in closed form:
This is the OLS estimator1. Every property in the following sections is a statement about this .
Error Assumptions and the Guaranteed Properties
The statistical properties of OLS results accumulate step by step as the assumptions on the errors get stronger. Exogeneity alone establishes the unbiasedness of the coefficients. Adding homoscedasticity and uncorrelatedness establishes the Gauss-Markov theorem and makes the standard errors reflect the actual variability of the coefficient estimates. Adding normality on top makes interval estimation exact in finite samples. Keeping track of which result depends on which assumption lets you judge how much of the results remains trustworthy when an assumption is in doubt.
Exogeneity and the Unbiasedness of the Coefficients
Under the exogeneity assumption , the OLS estimator is unbiased. Rewriting the estimator as , its expectation conditional on the predictors is , which equals the true coefficients. Each individual estimate deviates from the true coefficients by the error-driven term , but this deviation does not lean in any particular direction. Unbiasedness guarantees exactly this absence of lean and nothing more.
Exogeneity fails when a variable that affects the response and correlates with the included predictors is left out of the predictors, or when the relation between the response and the predictors is nonlinear. When a variable that affects the response and also correlates with the included predictors is missing from the model, its effect leaks into the coefficients of the correlated predictors. This distortion is called omitted variable bias. It is a distortion of the coefficients themselves, and collecting more observations does not remove it. When the conditional expectation of the response cannot be expressed as a linear combination of the predictors, the gap between the two remains in the error, and exogeneity fails in the same way. A failure of exogeneity can show up as a systematic pattern between the residuals and the fitted values, which the Residuals vs Fitted plot can reveal.
Homoscedasticity, Uncorrelatedness, and Standard Errors
The homoscedasticity and uncorrelatedness assumption requires the error variance to be constant across observations and the errors to be mutually uncorrelated. When this assumption holds in addition to exogeneity, the Gauss-Markov theorem states that the OLS estimator has the smallest variance among the unbiased estimators expressible as linear transformations of 2. The theorem does not require normality of the errors.
Homoscedasticity and uncorrelatedness are also the premise of the formula for the coefficient standard errors. Under this assumption, the coefficient variance simplifies from the general form to . The standard error replaces the in this formula with the estimate defined in the next section. When the error variance differs across observations or the errors are correlated, this simplification no longer holds. The coefficients remain unbiased under exogeneity alone, but the standard errors and the confidence intervals built on them stop reflecting the actual variability of the coefficient estimates. Unequal variance can be checked with the Scale-Location plot.
Normality and Interval Estimation
Assuming further that the errors are normal, , makes interval estimation based on the t distribution exact in finite samples. Under this assumption, follows a t distribution with degrees of freedom exactly. The probability that an interval contains the true coefficient is called the coverage probability, and the value set as the confidence level () is called the nominal level. The coverage probability of the confidence interval equals the nominal level regardless of the number of observations. The normality of the residuals can be checked with the Normal Q-Q plot.
The same assumption yields t-based intervals for the fitted values as well. Using the leverage , defined in Standardized Residuals and Diagnostic Statistics, the confidence interval (CI) for the mean response and the prediction interval (PI) for an individual observation are:
The and under the square roots correspond to the variances of the error between the fitted value and each interval's target. The target of the confidence interval is the mean response , and the estimation error has variance . The target of the prediction interval is a new observation at the same predictor values. Because the new observation's error is independent of the errors used to estimate , the prediction error has variance equal to the sum of the estimation error variance and the new error variance , which is .
The coverage probability of the coefficient confidence intervals approaches the nominal level in large samples even without normality. Suppose the errors are mutually independent draws from a common distribution with finite variance, and converges to a nonsingular matrix. Then by the central limit theorem, converges in distribution to a normal distribution. When it converges in distribution this way, is said to have asymptotic normality. Through this asymptotic normality, the coverage probability of the t-based confidence intervals approaches the nominal level as the sample grows.
The coverage probability of the prediction intervals does not approach the nominal level as the sample grows when normality fails. While the coefficient confidence interval concerns the distribution of the estimator , the prediction interval depends on the tails of the distribution of the new observation's error itself. This error distribution does not change as the sample grows, so the large-sample normal approximation is not available for prediction intervals.
Residual Degrees of Freedom and the Variability of the Error Variance Estimate
The error variance is estimated by : the residual sum of squares (RSS) divided by the residual degrees of freedom . Under exogeneity, homoscedasticity, and uncorrelatedness, is an unbiased estimator of . The residual standard error shown in the Linear Regression tab is its square root .
How much varies around the true value is determined by the residual degrees of freedom. Under normality, follows a chi-squared distribution with degrees of freedom, and the coefficient of variation of (its standard deviation divided by its mean) is . With 50 residual degrees of freedom, the coefficient of variation is 0.2, so the estimate varies around the true value by roughly 20 percent. With 2 residual degrees of freedom, the coefficient of variation reaches 1. From the chi-squared quantiles, at 2 residual degrees of freedom the 90% interval of is roughly . At this degrees of freedom, redrawing the data gives an error variance estimate between one twentieth and three times the true value with probability 0.9.
Every quantity that uses inherits this uncertainty. The widths of the standard errors, confidence intervals, and prediction intervals are proportional to , and the effect sizes in the analysis of variance (ANOVA) also take the error variance estimate as an input. When normality holds exactly, the t distribution accounts for the variation of , so the coverage probability of the confidence intervals stays at the nominal level. The interval widths, however, vary strongly from data to data in proportion to .
Apart from propagating the variation of , small residual degrees of freedom also enlarge the t quantiles. The 97.5% quantile at 2 degrees of freedom is 4.30, more than twice the normal distribution's 1.96. Furthermore, when normality fails, the large-sample approximation of the previous section that would compensate for it is not available at small degrees of freedom.
There are two ways to increase the residual degrees of freedom: collect more observations, or estimate fewer parameters. Candidates for removal are interaction terms whose absence of effect is theoretically certain, and predictors with no conceivable relation to the response. Removing them from the model frees up residual degrees of freedom and stabilizes the error variance estimate.
Standardized Residuals and Diagnostic Statistics
Because the errors are unobservable, diagnostics use the residuals , but residuals do not share the errors' properties. The residual vector can be written as , where is the hat matrix defined under Leverage, so even under exogeneity, homoscedasticity, and uncorrelatedness, : the variance of the raw residuals differs across observations. This section defines the standardized residuals that correct this unevenness and the influence measure derived from them. See the Linear Regression tab for how the diagnostic plots use these quantities.
Leverage
The leverage is a quantity that grows as observation 's predictor values sit farther out in the predictor space. The vector of fitted values can be written as , and this is called the hat matrix. The leverage is its diagonal element , which is the weight with which observation 's own value contributes to its own fitted value . In models with an intercept, it decomposes as , where is observation 's predictor values excluding the intercept with each predictor's mean subtracted, and stacks these vectors for all observations. The farther an observation lies from the predictor means, the larger its leverage. In models without an intercept, leverage measures the distance from the origin instead of the mean.
The possible values of leverage are constrained: and always hold. The average is therefore .
At a high-leverage observation, the fit depends strongly on that observation's own value. Since is the weight of in , an observation with close to 1 has a fitted value that is nearly itself, and its residual variance approaches 0. A poor fit at that position barely shows in the raw residual, so looking at residual sizes alone cannot detect this dependence. The standardized residuals that follow correct the unevenness of the residual variances, and Cook's Distance quantifies this dependence as an influence measure.
Standardized Residuals
A standardized residual divides the raw residual by an estimate of its standard deviation, putting the observations on a common scale:
By correcting the unevenness , standardized residuals have mean 0 and variance 1 at every observation when all the assumptions hold. This definition, which uses all observations to estimate the variance, is called the internally studentized residual3.
Cook's Distance
Cook's Distance is an influence measure that summarizes in a single number how much the coefficient estimates change when observation is removed and the model is refitted (Cook, 1977). Its definition rests on the difference between the coefficients before and after removal, but it can be rearranged into a form computed from the standardized residual and the leverage alone:
The influence is the product of the residual size and the leverage term . An observation with a zero residual () leaves the coefficients unchanged upon removal, however high its leverage. An observation near the center of the predictor space () barely moves the coefficients even when its residual is large. The leverage term, on the other hand, grows without bound as approaches 1, so an observation with leverage close to 1 can have a large even when its standardized residual is of ordinary size.
At an observation with leverage 1, neither the standardized residual nor Cook's Distance is defined: , so both denominators are 0. Leverage 1 means the fitted value equals the observed value () whatever is; the regression necessarily passes through this observation. For example, suppose some predictor takes a nonzero value only at this one observation. Its coefficient settles at whatever value reproduces this observation exactly, without affecting the fit of the others. As a result, this observation's residual is always 0 regardless of the value of , and the residual-based standardized residual and Cook's Distance carry no information. Being uncomputable does not mean the influence is small. It signals the most extreme form of dependence: part of the coefficients is determined by this single observation alone.
Multicollinearity and VIF
Multicollinearity is a state in which the predictors are strongly linearly correlated. It enlarges the variance of the coefficient estimators, that is, how much the coefficient estimates vary when the data are redrawn. Between correlated predictors, the data can barely distinguish which coefficient the response's variation should be attributed to, and the combination of coefficient estimates swings widely with small differences in the data.
The VIF (Variance Inflation Factor) measures how much of predictor 's variation remains usable for estimating the coefficient . Let be the coefficient of determination from regressing on the remaining predictors with an intercept; then . Since is the share of 's variation the other predictors can explain, is the share left unexplained.
What the estimation of can use is only the part of 's variation not tied to the other predictors: the residual from regressing on the remaining predictors. This is because the coefficient represents the change in the response when moves with the other predictors held fixed. The multiple-regression equals the coefficient of the simple regression on this residual alone, and in models with an intercept its variance is
The denominator is exactly the usable remainder of 's total variation . With , only a quarter of 's variation contributes to estimating , and the standard error is inversely proportional to the square root of this usable variation4.
What multicollinearity enlarges is the variance of the individual coefficient estimators; the model's overall predictions do not become unstable to the same degree. The fitted values are determined on the space spanned by the predictors, so assigning the coefficients to either of two correlated predictors changes the predictions very little. This is why the model's coefficient of determination and predictions can stay stable while only the individual coefficients and standard errors are unstable. How serious multicollinearity is depends on whether the goal is interpreting coefficients or making predictions.
In models without an intercept, the VIF as defined here does not apply, because the in its definition is the coefficient of determination of a regression with an intercept.
Multicollinearity degrades not only the statistical certainty of the coefficients but also the numerical accuracy of the computation. The stronger the correlation among the predictors, the larger the condition number of the design matrix, and the more the rounding errors are amplified. See Numerical Fundamentals for the mechanism and remedies.
See also
- Linear Regression tab - Operating OLS and reading the results screen and diagnostic plots
- GLM Fundamentals - The mathematics of generalized linear models, which include OLS as a special case
- Numerical Fundamentals - The condition number of the design matrix and rounding errors
- Glossary - Definitions of unbiasedness, consistency, asymptotic normality, and more
References
- Cook, R. D. (1977). Detection of influential observation in linear regression. Technometrics, 19(1), 15-18. https://www.jstor.org/stable/1268249
Footnotes
-
MIDAS does not compute the value of from this formula. It uses a QR decomposition of the design matrix instead, because forming squares the condition number and amplifies the effect of rounding errors. The computed result agrees with the formula above up to rounding error. See Numerical Fundamentals for the relation between the condition number and rounding errors. ↩
-
An estimator with this property is called the BLUE (Best Linear Unbiased Estimator). The theorem compares only linear unbiased estimators, and it makes no claim about the absolute size of the variance. ↩
-
There is also a definition that excludes observation itself from the variance estimate (the externally studentized residual). MIDAS's diagnostic plots and diagnostic datasets use the internally studentized values. ↩
-
The name "variance inflation factor" comes from reading the same formula as "how many times larger the variance is than , the variance if were uncorrelated with the other predictors ()." ↩
Also available as a Markdown file.