Skip to main navigation Skip to search Skip to main content

Comparison of Pipelines, Seq2seq Models, and LLMs for Rare Disease Information Extraction

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

End-to-end relation extraction (E2ERE) is an important application of natural language processing (NLP) in biomedicine. The extracted relations populate knowledge graphs and drive more high level applications in knowledge discovery and information retrieval. E2ERE is frequently handled at the sentence level involving continuous entities. A more complex setting is document level E2ERE with discontinuous and overlapping/nested entities. We identified a recently introduced RE dataset for rare diseases (RareDis) that has these complex traits. Among current E2ERE methods, we see three well-known paradigms: (1) pipeline based approaches where a named entity recognition (NER) model’s output is input to a relation classification (RC) model; (2) joint sequence-to-sequence style models where the raw input text is directly transformed into relations through linearization schemas; and (3) generative large language models (LLMs), where prompts, fine-tuning, and in-context learning are being leveraged for RE. While LLMs are becoming popular because of tools such as ChatGPT, the biomedical NLP community needs to carefully evaluate which paradigm is more suitable for E2ERE. In this effort, using the RareDis dataset as a complex use-case, we evaluate the best representative models from each of the three paradigms for E2ERE. Our findings reveal that pipeline models are still the best, while sequence-to-sequence models are not far behind. We verify these findings on a second E2ERE dataset for chemical-protein interactions. Although LLMs are more suitable for zero-shot settings, our results show that it is better to work with more conventional models trained and tailored for E2ERE when training data is available. Our contribution is also the first to conduct E2ERE for the RareDis dataset.

Original languageEnglish
Title of host publicationNatural Language Processing and Information Systems - 30th International Conference on Applications of Natural Language to Information Systems, NLDB 2025, Proceedings
EditorsRyutaro Ichise
Pages49-63
Number of pages15
Volume15836
DOIs
StatePublished - 2026
Event30th International Conference on Natural Language and Information Systems, NLDB 2025 - Kanazawa, Japan
Duration: Jul 4 2025Jul 6 2025

Publication series

NameLecture Notes in Computer Science
Volume15836 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference30th International Conference on Natural Language and Information Systems, NLDB 2025
Country/TerritoryJapan
CityKanazawa
Period7/4/257/6/25

Bibliographical note

Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.

Funding

FundersFunder number
NLM NIH HHSR01 LM013240

    Keywords

    • biomedical relation extraction
    • encoder-decoder models
    • end-to-end models
    • information extraction
    • large language models

    ASJC Scopus subject areas

    • Theoretical Computer Science
    • General Computer Science

    Fingerprint

    Dive into the research topics of 'Comparison of Pipelines, Seq2seq Models, and LLMs for Rare Disease Information Extraction'. Together they form a unique fingerprint.

    Cite this