Neural Networks and Diagnosis in the Clinical Laboratory: State of the Art
Author(s) -
Domenic V. Cicchetti
Publication year - 1992
Publication title -
clinical chemistry
Language(s) - English
Resource type - Journals
SCImago Journal Rank - 1.705
H-Index - 218
eISSN - 1530-8561
pISSN - 0009-9147
DOI - 10.1093/clinchem/38.1.9
Subject(s) - haven , state (computer science) , library science , art history , medicine , history , computer science , mathematics , algorithm , combinatorics
In the current issue of Clinical Chemistry, Astion and Wilding (1) compared quadratic discriniinant function analysis (QDFA) with the new technique of neural network (NN) modeling (2-4), for diagnosing the presence or absence of breast cancer. The authors are to be commended for presenting a lucid description of NN to this journal's readership as well as for discussing some of the limitations of their research design. Here I wish to discuss, in more detail, the further implications these limitations have for the design of future research involving NN and other multivariate techniques as approaches to clinical laboratory diagnosis. To accomplish this aim, I shall focus upon: (a) the Astion and Wilding investigation, (b) relevant findings from other investigations, and (c) some specific guidelines for further research in this important area of diagnostic inquiry. As Astion and Wilding note, a significant shortcoming of their study is the small number of patients in the training set relative to the number of predictor variables (nine predictor variables for NNs and seven for QDFA). As the subject-to-variable ratio decreases, the probability increases that one will observe a chance relationship between a predictor variable and an output category. These chance relationships tend to make the classification rate of the training group artificially high Moreover, they point out a secondproblem, the small size of the cross-validation set. The small cross-validation set makes it impossible to discern whether the difference between the cross-validation rates of the NN (80%) and the discriininant function (75%) is significant. In addition, the small cross-validation group decreases the accuracy of the shrinkage estimates (i.e., the decrease in classification rate when the method is applied to a different set of subjects from the set on which it was trained). The authors are quite correct. For example, the empirical work of Fletcher et al. (6) shows that one can expect, by chance alone, artificially high classification rates for linear discriminant function analysis (LDFA); these rates are directly dependent on the ratio of subjects to predictor variables, rather than on the number of subjects in each of the two groups. Thus, when the ratio is 1:1, whether the number of subjects in the two groups is 10,25, or 50, the expected shrinkage estimates vary within the narrow band of 34% to 36%. For a ratio between 9% and 12%. [It is interesting to note that increasing the ratio beyond 5:1 does not materially further reduce …
Accelerating Research
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom
Address
John Eccles HouseRobert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom