Abstract
Background: Almost all correlation measures currently available are unable to directly handle missing values. Typically, missing values are either ignored completely by removing them or are imputed and used in the calculation of the correlation coefficient. In either case, the correlation value will be impacted based on the perspective that the missing data represents no useful information. However, missing values occur in real datasets for a variety of reasons. In metabolomics datasets a major reason for missing values is that a specific measurable phenomenon falls below the detection limits of the analytical instrumentation (left-censored values). These missing data are not missing at random, but represent potentially useful information by virtue of their “missingness” at one end of the data distribution. Methods: To include this information due to left-censored missingness, we propose the information-content-informed Kendall-tau (ICI-Kt) methodology. We develop a statistical test and then show that most missing values in metabolomics datasets are the result of left-censorship. Next, we show how left-censored missing values can be included within the definition of the Kendall-tau correlation coefficient, and how that inclusion leads to an interpretation of information being added to the correlation. We also implement calculations for additional measures of theoretical maxima and pairwise completeness that add further layers of information interpretation in the methodology. Results: Using both simulated and over 700 experimental data sets from the Metabolomics Workbench, we demonstrate that the ICI-Kt methodology allows for the inclusion of left-censored missing data values as interpretable information, enabling both improved determination of outlier samples and improved feature–feature network construction. Conclusions: We provide explicitly parallel implementations in both R and Python that allow fast calculations of all the variables used when applying the ICI-Kt methodology on large numbers of samples. The ICI-Kt methods are available as an R package and Python module on GitHub.
| Original language | English |
|---|---|
| Article number | 245 |
| Journal | Metabolites |
| Volume | 16 |
| Issue number | 4 |
| DOIs | |
| State | Published - Apr 2026 |
Bibliographical note
Publisher Copyright:© 2026 by the authors.
Funding
This work was supported in part by the grants NSF 2020026 (PI Moseley), NIH 1R03LM014928-01 (PI Moseley), NSF ACI1626364 (Griffioen, Moseley), P30 CA177558 (PI Evers) via the Markey Cancer Center Biostatistics and Bioinformatics Shared Resource Facility (MCC BB-SRF), P20 GM121327 (PD St. Clair), and P42 ES007380 (PI Pennell) via the Data Management and Analysis Core (DMAC).
| Funders | Funder number |
|---|---|
| National Science Foundation Arctic Social Science Program | 2020026 |
| National Institutes of Health (NIH) | P30 CA177558, 1R03LM014928-01, ACI1626364 |
| Markey Cancer Center Biostatistics and Bioinformatics Shared Resource Facility | P20 GM121327, P42 ES007380 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 3 Good Health and Well-being
Keywords
- correlation
- left-censored
- metabolomics
- missingness
ASJC Scopus subject areas
- Endocrinology, Diabetes and Metabolism
- Biochemistry
- Molecular Biology
Fingerprint
Dive into the research topics of 'Information-Content-Informed Kendall-Tau Correlation Methodology: Interpreting Missing Values in Metabolomics as Potentially Useful Information'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver