Longitudinal Mixed Model Workflow

Overview

A longitudinal mixed model workflow is a practical way to build linear mixed models for repeated measurements over time.

Core logic

Where is the variation?
-> How does the outcome change over time?
-> How do individuals differ in that change?
-> What explains those differences?

Important

Build model complexity only when it corresponds to a meaningful scientific question.

Composite model form

For longitudinal data, level-1 and level-2 models can be combined into one equation.

For a random-intercept + random-slope model:

Yij=γ00+γ10TIMEij+ζ0i+ζ1iTIMEij+ϵij

The residual structure naturally induces dependence among repeated measurements from the same person.

Before modeling

Before fitting models, examine:

The goal is to see whether systematic change and meaningful individual heterogeneity are present.

Model-building strategy

A useful strategy is to build the model gradually:

Unconditional means→Unconditional growth→Conditional growth

Each new model is interpreted relative to a simpler baseline.

1. Unconditional means model

No TIME or substantive predictors:

Yij=γ00+ζ0i+ϵij

Purpose:

No systematic change over time is modeled yet.

2. Unconditional growth model

Add TIME:

Yij=γ00+γ10TIMEij+ζ0i+ζ1iTIMEij+ϵij

Purpose:

A significant random-slope variance suggests meaningful heterogeneity in rates of change that may be explained by predictors.

3. Conditional growth model

Add substantive predictors based on the research question.

Predictors can be:

Time-invariant predictors

A time-invariant predictor has one value for each individual.

Examples:

Time-invariant predictors can explain individual differences in intercepts and/or slopes.

For example:

π0i=γ00+γ01Xi+ζ0iπ1i=γ10+γ11Xi+ζ1i

To explain differences in rates of change, include an interaction with TIME.

Time-varying predictors

A time-varying predictor can take different values for the same individual at different measurement occasions.

Examples:

These predictors enter the level-1 model because they vary within individuals over time.

A simple form is:

Yij=π0i+π1iTIMEij+π2Xij+ϵij

where Xij is a time-varying predictor.

π2 describes the association between the outcome and the predictor at a particular occasion, conditional on the other terms in the model.

Centering time-varying predictors

Time-varying predictors can be represented in several ways, depending on the research question.

Possible representations include:

These representations can change the interpretation of model parameters even when the underlying information is similar.

For example, within-person centering uses:

Xij−X¯i

This represents how much the predictor at occasion j differs from that individual's own average.

This can help distinguish:

Tip

The choice of centering should be driven by substantive interpretation rather than a universal rule.

Endogeneity

Associations involving time-varying predictors can be difficult to interpret causally.

If Xij and Yij change together, it may be unclear whether X→Y or Y→X.

This is especially important when the individual's current outcome can influence the value of the predictor itself. See Causal inference.

Functional form of time

Time can be modeled in different ways:

Choose the simplest form that adequately represents the data and research question.

Tip

The meaning of the intercept depends heavily on how TIME is coded. If TIME=0 means baseline, the intercept usually represents expected baseline status.

Estimation & Model Comparison

Estimation

Mixed models can be estimated using methods such as:

Tip

GLS = extending ordinary least squares by allowing residuals to be correlated and heteroscedastic, which is important for longitudinal data.

Comparing models

For models estimated by ML, model fit can be compared using deviance:

D=−2log⁡L

Smaller deviance indicates better fit.

For two nested models:

ΔD=Dreduced−Dfull

This can be tested approximately using a χ2 distribution, with degrees of freedom equal to the difference in the number of estimated parameters.

This is the logic of the Likelihood ratio test.

Important

Use ML, not REML, when comparing models that differ in fixed effects. REML is often preferred for estimating variance components in a final model.

Variance reduction and pseudo-R2

Variance components from simpler models can serve as baselines for evaluating later models.

For example:

Rpseudo2=σbaseline2−σnew2σbaseline2

This describes the proportional reduction in unexplained variance after adding predictors.

Different variance components can have different pseudo-R2 values:

Interpretation:

Warning

Pseudo-R2 should be interpreted cautiously because variance components can occasionally increase after adding predictors.

Model assumptions and diagnostics

Important assumptions should be examined at both levels.

Functional form

Residuals

Inspect residuals for:

Diagnostics are especially important for the final model being interpreted.

Interpretation

Fixed effects

Fixed effects describe the systematic population-level trajectory and predictor effects.

Examples:

Random effects

Random effects describe how individual trajectories deviate from the population-average trajectory.

Important variance components:

Hint

Variance components represent remaining unexplained heterogeneity conditional on the current fixed effects.

Empirical Bayes estimates

Mixed models can estimate individual trajectories using model-based / empirical Bayes estimates.

These combine:

population information+individual information

rather than estimating each person's trajectory independently with OLS.

The resulting individual estimates are usually shrunk toward the relevant population-average trajectory, especially when an individual's data are sparse or noisy.

This often improves precision, but the quality of empirical Bayes estimates depends on the correctness of the fitted model.

Practical workflow

Core

First quantify variation, then model change over time, then explain individual differences in that change.

  1. Explore the longitudinal structure.
  2. Fit an unconditional means model.
  3. Add TIME to fit an unconditional growth model.
  4. Choose the functional form of TIME.
  5. Add substantive predictors.
  6. Decide how predictors should be centered or represented.
  7. Add interactions with TIME if the question concerns change.
  8. Examine variance reduction.
  9. Compare nested models when appropriate.
  10. Check assumptions and model fit.
  11. Interpret fixed effects, random effects, and variance components.