Author's School

School of Medicine

ORCID

https://orcid.org/0000-0001-7525-6471

Author's Department/Program

Biology and Biomedical Sciences

Language

English (en)

Date of Award

12-15-2026

Degree Type

Dissertation

Degree Name

Doctor of Philosophy (PhD)

Chair and Committee

Lei Liu

Committee Members

Yin Cao, Jun Chen, Gautam Dantas, Kristine M. Wylie

Abstract

Differential abundance analysis in microbiome studies aims to identify taxa whose abundance differs across biological or clinical conditions. The observed data are typically taxon-specific sequencing read counts, representing reads assigned to different taxa within each sample. These counts are indirect measurements of the underlying microbial abundance profile and are constrained by sample-specific library sizes. Microbiome count data are also typically sparse, overdispersed, and heteroscedastic. Together, these characteristics create substantial challenges for differential abundance analysis and make the results highly sensitive to normalization procedures, model specification, and the statistical methods used for inference.

Normalization defines the scale on which samples are compared, so a biased size factor can place the estimated effect on a distorted scale. In regression-based analysis, the mean model defines the taxon-level effect estimated on that scale, often expressed as a log fold change for the condition or covariate of interest. The variance model specifies the working mean-variance relationship. When this relationship is misspecified, model-based standard errors may be invalid. Consequently, the choice of standard-error estimator can affect the resulting p-values and, ultimately, which taxa are identified as differentially abundant.

Our first study addresses model specification for heteroscedastic microbiome counts. We develop a flexible quasi-likelihood model that estimates the mean-variance relationship as a smooth function of the mean. The method does not require a fully specified count distribution. It retains regression-based effect estimation while reducing dependence on a prespecified parametric variance function.

The second study addresses standard-error estimation under variance misspecification. It evaluates bootstrap and sandwich standard-error estimators for count-based regression. A Poisson model is used as a working mean model, and empirical covariance estimators replace the model-based covariance. This separates estimation of the mean effect from estimation of its standard error and evaluates whether statistical inference remains reliable when the working variance model is misspecified.

The third study addresses normalization in the presence of compositional bias. We develop Iterative Reference Selection (IRS), a reference-based normalization method for microbiome differential abundance analysis. IRS uses the group or covariate of interest to refine the reference set and exclude taxa likely to be differentially abundant. The selected reference set is then used to estimate sample-specific size factors. This reduces reference contamination and makes downstream modeling less sensitive to compositional bias.

Together, these studies examine related statistical challenges in microbiome differential-abundance analysis: normalization error, mean-variance misspecification, and standard-error estimation under variance misspecification. The methods developed here connect normalization, modeling, and inference through their effects on regression-based differential-abundance evidence.

DOI

https://doi.org/10.48765/sjds-t773

Share

COinS