Regression

Summary
Related Notes

Regression Models Taxonomy

Understanding Regression Model Family

Regression models can be viewed as combinations of three independent dimensions:

Regression Model Family

Model Response distribution Functional Form Random Effects Core Formula Typical Use
Linear Model (LM) Gaussian Linear ❌ Y=Xβ+ε Standard linear regression
Generalized Linear Model (GLM) Any exponential family Linear ❌ g(μ)=Xβ Logistic, Poisson, Gamma regression
Generalized Additive Model (GAM) Gaussian (or others) Nonlinear ❌ g(μ)=β0+∑jfj(Xj) Flexible nonlinear effects
Linear Mixed Model (LMM) Gaussian Linear ✅ Y=Xβ+Zb+ε Continuous longitudinal / hierarchical data
Generalized Linear Mixed Model (GLMM) Non-Gaussian Linear ✅ g(μ)=Xβ+Zb Binary or count longitudinal data
Generalized Additive Mixed Model (GAMM) Any Nonlinear ✅ g(μ)=β0+∑jfj(Xj)+Zb Longitudinal data with nonlinear trajectories
For formulas:

  • Xβ = fixed effects
  • Zb = random effects
  • g(⋅) = link function (identity, logit, log, ...)
  • f(⋅) = smooth function (typically splines)
    ​

Relationship Between Models

LM -> GLM

GLM extends LM by allowing non-Gaussian responses.

GLM -> GAM

GAM extends GLM by replacing linear terms with smooth functions: instead of β×X, use f(X)

LM -> LMM

LMM extends LM by adding random effects.

GLM + LMM -> GLMM

GLMM = GLM + Random Effects ("mixed")

GAM + LMM -> GAMM

GAMM = GAM + Random Effects ("mixed")

Rule of Thumb

%%{init: {"themeVariables": {"fontSize": "12px"}, "flowchart": {"nodeSpacing": 25, "rankSpacing": 30}}}%%
flowchart TD
    A[Choose model] --> B{Repeated / clustered?}

    B -->|Yes| C{Nonlinear effects?}
    B -->|No| D{Nonlinear effects?}

    C -->|Yes| GAMM[GAMM]
    C -->|No| E{Outcome?}

    D -->|Yes| GAM[GAM]
    D -->|No| F{Outcome?}

    E -->|Continuous| LMM[LMM]
    E -->|Binary| GLMM1[Logistic GLMM]
    E -->|Count| GLMM2[Poisson GLMM]

    F -->|Continuous| LM[LM]
    F -->|Binary| GLM1[Logistic GLM]
    F -->|Count| GLM2[Poisson GLM]

Regression Methods and Extensions

Core regression models by outcome type

Linear Regression

Logistic Regression

Poisson Regression

Regression with transformed predictors

Principal Component Regression

Source

  1. apply PCA (Dimensionality Reduction#Principal Component Analysis (PCA)) to generate principal components from the predictor variables, with the number of principal components matching the number of original features p
  2. keep the first k principal components that explain most of the variance (where k < p), where k is determined by cross-validation
  3. fit a linear regression model on these k principal components

Partial Least Squares Regression

Regularized regression

Regularized regression adds a penalty (Regularization) term to the loss function to control model complexity and reduce overfitting (Cost Functions#Cost function with regularization) .

Lasso Regression

Ridge Regression

Elastic Net

Regression Diagnostics

Linear regression diagnostics

residuals vs fitted

Q-Q plot

scale-location plot

leverage / Cook's distance

multicollinearity / VIF

R-Squared: coefficient of determination

R2=1−residual sum of squares (RSS)total sum of squares (TSS)=1−∑(yi−yi^)2∑(yi−y―)2

You can regard the R-Squared as how much the total variance of y is captured by the model (rather than errors), i.e.

var(y)=var(Xβ^)+var(e)R2=var(Xβ^)var(y)

so it measures the goodness of fit, but it does not validate the model.

Adjusted R-Squared

Adjusted R-squared is a modified version of R-squared that has been adjusted for the number of predictors in the model.

Adj R2=1−(1−R2)n−1n−p−1

where p is the number of regressors/predictors, n is the sample size.

Why you need adjusted R-squared?

It is important, because adding independent variables will make the R-squared never decrease, then you cannot tell whether the increase of R-squared is due to the goodness of fit or more variables.
It includes the penalising factor that penalises you for adding independent variables that don't help your model.

GLM and classification diagnostics