Course Text: Pages 7-9, 20-21, C1: Pages 8-9, C2: Pages 19-21
Several different criteria have been proposed to assist in determining the number of classes –that is, for deciding between one candidate model that hypothesizes say 3 latent classes over another postulating say 4 classes. Since the various criteria are each justified under different assumptions, it should not be a surprise that they do not always agree with each other.
LC modeling was first introduced for analyzing a relatively small number of categorical indicators and the p-value associated with the likelihood ratio (L2) and Pearson (X2) chi-squared statistics were used to determine the number of classes. However, modern applications often involve many indicators, among which one or more continuous variables may also be included (to be discussed in Topic J), which yield sparse data. In these cases, the p-value obtained from the chi-squared statistics is often not valid in practice (due to the sparse data). As such, information criteria have replaced the p-value as the criteria most used in current applications.
While the literature is mixed on which of several information criteria is preferable, the Bayesian Information Criterion (BIC) is the one that is most often reported in practice. The BIC and related information criteria are provided as standard LatentGOLD output.
More recently, two other approaches have been gaining popularity in determining the number of classes :
C1. Bootstrap and conditional bootstrap
(LG tutorial 1: pages 18-20)
Since the chi-squared assumption used to compute the p-value obtained from L2 is violated when the data is sparse, the p-value reported in LatentGOLD is often not valid in practice. An alternative is to estimate the p-value without relying on the chi-squared assumption, using the bootstrap. In addition, the conditional bootstrap can be used to choose between two nested models, say to choose between 3 and 4 class models. LG tutorial 1 illustrates how both of these approaches are performed with LatentGOLD.
C2. Latent Class Tree Modeling
(Course Text: pages 20-21, Latent Class Tree Tutorial)
Unlike the approaches above that utilize statistical significance to determine the number of classes, the LC Tree modeling is a radical alternative which prioritizes substantive significance over statistical significance, resulting in latent classes that are more interpretable. Specifically, the more traditional information criteria for determining the number of classes in exploratory settings continues to increase the number of classes so long as the BIC statistic keeps declining. The model with the lowest BIC value is then retained as the final model. Proponents of LC Tree modeling point out two problems with this approach:
1) it may yield a final model that is difficult to interpret (especially when many indicators are included in the model), in part because it contains too many latent classes, and
2) it is difficult to compare various solutions with different numbers of classes to see how they are related, and in particular, solutions with fewer classes than that suggested by BIC may also be of interest.
To overcome these two problems, Van den Bergh, Vermunt, and colleagues proposed a new LC modeling approach called latent class tree (LCT) modeling, which is similar to (divisive) hierarchical cluster analysis. We will learn more about LC Tree modeling in Topic F, where you will have the opportunity to try out new interactive features that have been implemented in LatentGOLD 6.0.
Assigned Reading:
LG tutorial 1:
C1: Bootstrap, (pages 7-11)
C2: Conditional Bootstrap, (pages 19-21)