Computing Identity Co-Reference Across Drug Discovery Datasets

Brenninkmeijer, Christian Y. A.; Dunlop, Ian; Goble, Carole; Gray, Alasdair J. G.; Pettifer, Steve; Stevens, Robert

Tools

Export citation

Search in Google Scholar

Computing Identity Co-Reference Across Drug Discovery Datasets

Proceedings article published in 2013 by Christian Y. A. Brenninkmeijer, Ian Dunlop, Carole Goble

, Alasdair J. G. Gray, Steve Pettifer, Robert Stevens

This paper is available in a repository.

Full text: Download

Preprint: policy unknown

Upload

Postprint: policy unknown

Upload

Published version: policy unknown

Upload

Abstract

This paper presents the rules used within the Open PHACTS (http://www.openphacts.org) Identity Management Service to compute co-reference chains across multiple datasets. The web of (linked) data has encouraged a proliferation of identifiers for the concepts captured in datasets; with each dataset using their own identifier. A key data integration challenge is linking the co-referent identifiers, i.e. identifying and linking the equivalent concept in every dataset. Exacerbating this challenge, the datasets model the data differently, so when is one representation truly the same as another? Finally, different users have their own task and domain specific notions of equivalence that are driven by their operational knowledge. Consumers of the data need to be able to choose the notion of operational equivalence to be applied for the con- text of their application. We highlight the challenges of automatically computing co-reference and the need for capturing the context of the equivalence. This context is then used to control the co-reference computation. Ultimately, the context will enable data consumers to decide which co-references to include in their applications.

Links

Tools

Computing Identity Co-Reference Across Drug Discovery Datasets

Abstract