Scinovex
articleTop 10% cited

The Calculation of Posterior Distributions by Data Augmentation

Journal of the American Statistical Association · 1987 · Vol. 82(398) · pp. 528–540
Martin A. TannerWing Hung Wong

Abstract

Abstract The idea of data augmentation arises naturally in missing value problems, as exemplified by the standard ways of filling in missing cells in balanced two-way tables. Thus data augmentation refers to a scheme of augmenting the observed data so as to make it more easy to analyze. This device is used to great advantage by the EM algorithm (Dempster, Laird, and Rubin 1977) in solving maximum likelihood problems. In situations when the likelihood cannot be approximated closely by the normal likelihood, maximum likelihood estimates and the associated standard errors cannot be relied upon to make valid inferential statements. From the Bayesian point of view, one must now calculate the posterior distribution of parameters of interest. If data augmentation can be used in the calculation of the maximum likelihood estimate, then in the same cases one ought to be able to use it in the computation of the posterior distribution. It is the purpose of this article to explain how this can be done. The basic idea is quite simple. The observed data y is augmented by the quantity z, which is referred to as the latent data. It is assumed that if y and z are both known, then the problem is straightforward to analyze, that is, the augmented data posterior p(θ | y, z) can be calculated. But the posterior density that we want is p(θ | y), which may be difficult to calculate directly. If, however, one can generate multiple values of z from the predictive distribution p(z | y) (i.e., multiple imputations of z), thenp(θ | y) can be approximately obtained as the average ofp(θ | y, z) over the imputed z's. However, p(z | y) depends, in turn, onp(θ | y). Hence if p(θ | y) was known, it could be used to calculate p(z | y). This mutual dependency between p(θ | y) and p(z | y) leads to an iterative algorithm to calculate p(θ | y). Analytically, this algorithm is essentially the method of successive substitution for solving an operator fixed point equation. We exploit this fact to prove convergence under mild regularity conditions. Typically, to implement the algorithm, one must be able to sample from two distributions, namely p(θ | y, z) andp(z | θ, y). In many cases, it is straightforward to sample from either distribution. In general, though, either sampling can be difficult, just as either the E or the M step can be difficult to implement in the EM algorithm. For p(θ | y, z) arising from parametric submodels of the multinomial, we develop a primitive but generally applicable way to approximately sample θ. The idea is first to sample from the posterior distribution of the cell probabilities and then to project to the parametric surface that is specified by the submodel, giving more weight to those observations lying closer to the surface. This procedure should cover many of the common models for categorical data. There are several examples given in this article. First, the algorithm is introduced and motivated in the context of a genetic linkage example. Second, we apply this algorithm to an example of inference from incomplete data regarding the correlation coefficient of the bivariate normal distribution. It is seen that the algorithm recovers the bimodal nature of the posterior distribution. Finally, the algorithm is used in the analysis of the traditional latent-class model as applied to data from the General Social Survey.

Statistical Methods and Bayesian InferenceBayesian Methods and Mixture ModelsOptimal Experimental Design MethodsPosterior probabilityBayesian probabilityMathematicsSimple (philosophy)Posterior predictive distributionApproximate Bayesian computationLikelihood functionMissing dataAlgorithmStatistics
Citations
3,753
FWCI
9.36
field-weighted impact
References
20
Percentile
99%
vs. same field & year
Citations per year
Cited by
Missing Data Analysis: Making It Work in the Real World
Annual Review of Psychology · 2008 · 5,983 citations
Approximate Bayesian Inference with the Weighted Likelihood Bootstrap
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 1994 · 1,509 citations
Factorial Hidden Markov Models
Machine Learning · 1997 · 1,178 citations
Bayesian Computation Via the Gibbs Sampler and Related Markov Chain Monte Carlo Methods
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 1993 · 1,718 citations
Sampling-Based Approaches to Calculating Marginal Densities
Journal of the American Statistical Association · 1990 · 6,639 citations
Estimation of Finite Mixture Distributions Through Bayesian Sampling
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 1994 · 908 citations
An Introduction to MCMC for Machine Learning
Machine Learning · 2003 · 2,387 citations
References
Finding the Observed Information Matrix When Using the <i>EM</i> Algorithm
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 1982 · 2,218 citations
Maximum Likelihood from Incomplete Data Via the <i>EM</i> Algorithm
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 1977 · 49,286 citations
Bayesian Inference in Statistical Analysis.
Journal of the American Statistical Association · 1975 · 3,873 citations
Linear Operators. Part I: General Theory.
American Mathematical Monthly · 1960 · 2,794 citations
Linear Statistical Inference and Its Applications.
Biometrics · 1975 · 3,620 citations
Journal of the Royal Statistical Society.
The Economic Journal · 1923 · 3,267 citations
Related articles
Approximate Bayesian Computation
PLoS Computational Biology · 2013 · 663 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.