z-logo
open-access-imgOpen Access
Significance testing for canonical correlation analysis in high dimensions
Author(s) -
Ian W. McKeague,
Xin Zhang
Publication year - 2021
Publication title -
biometrika
Language(s) - English
Resource type - Journals
SCImago Journal Rank - 3.307
H-Index - 122
eISSN - 1464-3510
pISSN - 0006-3444
DOI - 10.1093/biomet/asab059
Subject(s) - mathematics , canonical correlation , estimator , statistical hypothesis testing , statistics , distance correlation , random variable
We consider the problem of testing for the presence of linear relationships between large sets of random variables based on a post-selection inference approach to canonical correlation analysis. The challenge is to adjust for the selection of subsets of variables having linear combinations with maximal sample correlation. To this end, we construct a stabilized one-step estimator of the euclidean-norm of the canonical correlations maximized over subsets of variables of pre-specified cardinality. This estimator is shown to be consistent for its target parameter and asymptotically normal, provided the dimensions of the variables do not grow too quickly with sample size. We also develop a greedy search algorithm to accurately compute the estimator, leading to a computationally tractable omnibus test for the global null hypothesis that there are no linear relationships between any subsets of variables having the pre-specified cardinality. We further develop a confidence interval that takes the variable selection into account.

The content you want is available to Zendy users.

Already have an account? Click here to sign in.
Having issues? You can contact us here