Probabilistic record linkage
Author(s) -
Adrian Sayers,
Yoav Ben–Shlomo,
Ashley Blom,
Fiona Steele
Publication year - 2015
Publication title -
international journal of epidemiology
Language(s) - English
Resource type - Journals
SCImago Journal Rank - 3.406
H-Index - 208
eISSN - 1464-3685
pISSN - 0300-5771
DOI - 10.1093/ije/dyv322
Subject(s) - probabilistic logic , linkage (software) , record linkage , computer science , bayes' theorem , process (computing) , data mining , bayesian probability , artificial intelligence , population , genetics , medicine , biology , environmental health , gene , operating system
Studies involving the use of probabilistic record linkage are becoming increasingly common. However, the methods underpinning probabilistic record linkage are not widely taught or understood, and therefore these studies can appear to be a 'black box' research tool. In this article, we aim to describe the process of probabilistic record linkage through a simple exemplar. We first introduce the concept of deterministic linkage and contrast this with probabilistic linkage. We illustrate each step of the process using a simple exemplar and describe the data structure required to perform a probabilistic linkage. We describe the process of calculating and interpreting matched weights and how to convert matched weights into posterior probabilities of a match using Bayes theorem. We conclude this article with a brief discussion of some of the computational demands of record linkage, how you might assess the quality of your linkage algorithm, and how epidemiologists can maximize the value of their record-linked research using robust record linkage methods.
Accelerating Research
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom
Address
John Eccles HouseRobert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom