In the vast realm of healthcare, precise and accurate semantic translations are of paramount importance. The Unified Medical Language System (UMLS) has long been considered a valuable resource, offering tools and resources to aid in organizing and integrating diverse biomedical vocabularies. While UMLS holds value for semantic translations, it is essential to explore its pros and cons, particularly when it comes to synonymy.
A history of interoperability in healthcare
When the United States went on a journey of improving interoperability in healthcare, we were still exchanging information, dare I say knowledge, via snail mail and fax machines. Although we had computers, I can still remember the days of clinicians dictating notes into a Dictaphone, which assistants (me) then transcribed into a Word document stored on a local computer, printed out for signature, and placed into a paper chart. This then might be faxed to another clinician’s office for a referral or even mailed to a payer for precertification or in response to a request for records. Now that I have properly aged myself, let’s talk about what happens today. As they say with age comes wisdom and in the realm of healthcare interoperability that is somewhat true. Admittedly there are still faxes flying around out there but we are seeing a significant shift away from faxing, in particular by larger healthcare organizations.
We now exchange more healthcare data via electronic means than at any other time in history thanks to regulations by the ONC and CMS and hard work on the part of many non-governmental organizations such as HL7. In fact, according to the ONC’s latest report to Congress, in 2021 98% of hospitals use an ONC-certified EHR. Compare that to 28% in 2011. Additionally, 4 out of 5 office-based physician practices use an ONC-certified EHR. Compare that to 34% in 2011. One more important stat- nearly 2/3 of those clinicians utilize the services of a local or regional Health Information Exchange that allows them to exchange data with providers outside of their organizations. The creation of TEFCA in the 21st-Century Cures Act has accelerated this process and is providing a common framework that exchange.
Accurately coding clinical data is critical for achieving semantic interoperability
Given the adoption of technology, you may expect that you can log in to a patient portal and understand everything about your health care, and more importantly, if you land in the emergency room and are unable to speak for yourself, the clinician treating you would also have access to all pertinent information. Unfortunately, we aren’t there yet. The fact is that while we can share information on some level, we have a long way to go before that data is all in the same format and has the same meaning when it is received as when it is sent. I am excited to see all of this change because of FHIR but I digress.
The problem is not that we cannot share data - we certainly can, well mostly anyway. The problem is that the data needs to be understandable. This requires us to go beyond syntactic interoperability (the ability to exchange data electronically) and move into the realm of semantic interoperability. Many health IT professionals are very familiar with these data challenges:
- Structured and codified data (e.g., ICD-10-CM codes) is often not an accurate or complete representation of the patient’s medical conditions and status. It may be incomplete, codified to retired codes, out of date, and only tells a part of the patient’s story.
- Structured data that is not codified to standard terminology at all leaves the interpretation of the data up to the receiver. When a sender sends the acronym OM in a diagnosis field, do they mean otitis media or osteomyelitis?
- 80% of actionable Healthcare data is found in free text and images and needs to be interpreted by a clinician. It is not computer-readable and cannot be used in analytics.
As an industry, we rely on codified data for multiple use cases:
- Clinical Decision support needs fully codified data to accurately trigger algorithms and alerts
Population health initiatives such as reducing hospital re-admission rates or identifying gaps in care - Quality Reporting for CMS quality measures
- Communication between payers and providers
- True semantic interoperability
- Clinical research
Terminology server vs. UMLS: Which provides a higher level of accuracy for clinical data analytics?
I mentioned earlier that the UMLS is a valuable resource, and indeed with over 180 biomedical terminologies, it provides a rich library of terms that can be used to describe many clinical entities. Synonyms and relationships between the various terminologies create the potential to disambiguate much of the data in healthcare. However, any user of the UMLS needs to be aware that true synonymy is difficult to attain, especially with a large number of vocabularies, all created for different uses and different approaches to what terms are considered synonymous.
One of those terminologies is the Medical Subject Headings or MeSH. The MeSH is a controlled terminology designed to support the cataloging of biomedical literature, not the retrieval of information from electronic medical records or the clinical analysis of data. The UMLS, therefore, includes terms within MeSH concepts that may not be true synonyms. The MeSH exerts significant influence over the structure of the UMLS. This and nonsynonymous terms in the same concept brought in by other terminologies create the need to significantly modify UMLS content before it can be used in EHRs and other health IT applications.
Data captured at the point of care flows through to downstream analytics platforms that assist in population health analytics, closing gaps in care for HEDIS, and assessing cohorts for adherence to clinical quality measures. In those cases, analysts and data scientists may wish to query data sources for specific granular concepts. Nonsynonymy in UMLS concepts would lead to inaccurate information retrievals and noise. Conversely, leveraging a terminology server, you can easily group specific concepts into defined value sets if a broader search strategy is needed. The value sets can be built at various levels of granularity depending on the specific use case they are designed to address.
Regardless of your use case, the phrase, “garbage in, garbage out” remains true. You must always input clean, accurate, and reliable data in order to generate consistent, quality, and actionable analytics.