Course Text: Pages 24-25
In order to obtain a better understanding of the latent classes, it is often desirable to profile the latent class segments in terms of demographics or other exogenous variables. These exogeneous variables are referred to as covariates in the course text, and denoted Z1, Z2, …. . LatentGOLD allows covariates to be included in Step 1 in the LC modeling process in an active or inactive manner. Covariates are distinct from the indicators, which are denoted Y1, Y2, For the technical differences between variables selected as indicators and those selected as active covariates, see the course text.
When active or inactive covariates are included in a LatentGOLD run, additional output is produced. The LC Cluster solution with active covariates will generally be somewhat different from the solution without covariates. Alternatively, the “inactive covariates method” of entering covariates in a LC model, simply involves computing descriptive measures for the association between the covariates and the latent variable after estimating a model without covariates, and displaying the relationship between the covariates and the classes in tabular form.
Specifying covariates as active causes additional log-linear parameters to be included in the LC model (gammas), and estimated simultaneously with the other parameters (betas) and hence affect (somewhat) these model parameters. Like the other model parameters (betas), statistical tests are available for the gammas for assessing how well the covariates predict class membership. Because of the simultaneous estimation, the parameters from the LC model may differ from the same LC model without covariates. In some cases, the difference may be large, especially when several covariates are included in the model.
Inclusion of covariates in Step 3
While the ability to include covariates in LC modeling (Dayton & Macready, 1988) is an important extension of the traditional LC model, it is important to note that use of active covariates in the estimation step (Step 1) of LC modeling goes against the logic of most applied researchers, who view introducing covariates as a step that comes after the classification step (Step 2).
Additional problems with the inclusion of active covariates in Step 1 were pointed out by Vermunt (2010):
“However, the one-step approach has certain disadvantages. The first is that it may sometimes be impractical, especially when the number of potential covariates is large, as will typ¬ically be the case in a more exploratory study. Each time that a covariate is added or removed not only the predic¬tion model but also the measurement model needs to be re-estimated. A second disadvantage is that it introduces addi¬tional model building problems, such as whether one should decide about the number of classes in a model with or with¬out covariates. Third, the simultaneous approach does not fit with the logic of most applied researchers, who view intro¬ducing covariates as a step that comes after the classification model has been built. Fourth, it assumes that the classifica¬tion model is built in the same stage of a study as the model used to predict the class membership, which is not necessar-ily the case. It can even be that the researcher who constructs the typology using an LC model is not the same as the one who uses the typology in a next stage of the study (p. 451).
In addition, from a practical perspective, the class sizes may be quite different with and without covariates included in Step 1, especially when many covariates are included in the model. In this situation, a careful evaluation of these changes often reveals that the change in class size is an artifact caused by a violation of the assumption that the covariates are conditionally independent of indicators given the latent classes. This violation is known more generally as item bias or differential item functioning.
For all of these reasons, the bias-adjusted three-step approach to LC modeling (Bolck, Croon and Hagenaars, 2004; Vermunt, 2010), has become widely popular. As described on pages 24-25 of the Course Text, in the three-step approach, the LC model is estimated without covariates (step 1), respondents are assigned to classes (step 2), and the covariate effects are then estimated using the assigned class memberships (step 3). A generalized bias-adjusted step-three method (called ‘BCH’) as well as a conceptually simpler maximum likelihood approach have been implemented in the LatentGOLD Step3 module.
For further details, see section 6.1 of LatentGOLD Technical Guide,
Include covariates in Step 1 or Step 3?
The issue of whether to include covariates in Step 1 or Step 3 is not a closed issue. Both alternatives can be justified. The most recent publication on this topic (Vermunt and Magidson, 2020) points out that when direct effects exist between covariates and indicators, the situation is further complicated: See:
Structural Equation Modeling (Vermunt and Magidson, 2020); Interactive pdf
This issue will be discussed further in Topic K: Direct Effects, and in Exercise K1.
Assigned Reading:
Text (pages 24-25)
Updated Sage Article (updated in 2019):
H1: Multi-group Models, Section 3.3
H2: Covariates, Section 3.4
H3: Three-step LC Analysis, Section 3.5 (updated in 2019)
LatentGOLD Technical Guide
H3: LatentGOLD Technical Guide Section 3.3 (pages 24-25)
H4: Structural Equation Modeling (Vermunt and Magidson, 2020);
H5: Cambridge University Press :
Vermunt, J.K., and Magidson, J. (2002). Latent class cluster analysis. In: J. A. Hagenaars and A. L. McCutcheon (Eds.), Applied Latent Class Analysis, 89-106. Cambridge: Cambridge University Press.
H5: section on covariates (pages 5-6)