Applied Sciences (Jul 2023)

Semi-Automated Mapping of German Study Data Concepts to an English Common Data Model

  • Anna Chechulina,
  • Jasmin Carus,
  • Philipp Breitfeld,
  • Christopher Gundler,
  • Hanna Hees,
  • Raphael Twerenbold,
  • Stefan Blankenberg,
  • Frank Ückert,
  • Sylvia Nürnberg

DOI
https://doi.org/10.3390/app13148159
Journal volume & issue
Vol. 13, no. 14
p. 8159

Abstract

Read online

The standardization of data from medical studies and hospital information systems to a common data model such as the Observational Medical Outcomes Partnership (OMOP) model can help make large datasets available for analysis using artificial intelligence approaches. Commonly, automatic mapping without intervention from domain experts delivers poor results. Further challenges arise from the need for translation of non-English medical data. Here, we report the establishment of a mapping approach which automatically translates German data variable names into English and suggests OMOP concepts. The approach was set up using study data from the Hamburg City Health Study. It was evaluated against the current standard, refined, and tested on a separate dataset. Furthermore, different types of graphical user interfaces for the selection of suggested OMOP concepts were created and assessed. Compared to the current standard our approach performs slightly better. Its main advantage lies in the automatic processing of German phrases into English OMOP concept suggestions, operating without the need for human intervention. Challenges still lie in the adequate translation of nonstandard expressions, as well as in the resolution of abbreviations into long names.

Keywords