TY - GEN
T1 - Comparison of Pipelines, Seq2seq Models, and LLMs for Rare Disease Information Extraction
AU - Gupta, Shashank
AU - Ai, Xuguang
AU - Jiang, Yuhang
AU - Kavuluru, Ramakanth
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.
PY - 2026
Y1 - 2026
N2 - End-to-end relation extraction (E2ERE) is an important application of natural language processing (NLP) in biomedicine. The extracted relations populate knowledge graphs and drive more high level applications in knowledge discovery and information retrieval. E2ERE is frequently handled at the sentence level involving continuous entities. A more complex setting is document level E2ERE with discontinuous and overlapping/nested entities. We identified a recently introduced RE dataset for rare diseases (RareDis) that has these complex traits. Among current E2ERE methods, we see three well-known paradigms: (1) pipeline based approaches where a named entity recognition (NER) model’s output is input to a relation classification (RC) model; (2) joint sequence-to-sequence style models where the raw input text is directly transformed into relations through linearization schemas; and (3) generative large language models (LLMs), where prompts, fine-tuning, and in-context learning are being leveraged for RE. While LLMs are becoming popular because of tools such as ChatGPT, the biomedical NLP community needs to carefully evaluate which paradigm is more suitable for E2ERE. In this effort, using the RareDis dataset as a complex use-case, we evaluate the best representative models from each of the three paradigms for E2ERE. Our findings reveal that pipeline models are still the best, while sequence-to-sequence models are not far behind. We verify these findings on a second E2ERE dataset for chemical-protein interactions. Although LLMs are more suitable for zero-shot settings, our results show that it is better to work with more conventional models trained and tailored for E2ERE when training data is available. Our contribution is also the first to conduct E2ERE for the RareDis dataset.
AB - End-to-end relation extraction (E2ERE) is an important application of natural language processing (NLP) in biomedicine. The extracted relations populate knowledge graphs and drive more high level applications in knowledge discovery and information retrieval. E2ERE is frequently handled at the sentence level involving continuous entities. A more complex setting is document level E2ERE with discontinuous and overlapping/nested entities. We identified a recently introduced RE dataset for rare diseases (RareDis) that has these complex traits. Among current E2ERE methods, we see three well-known paradigms: (1) pipeline based approaches where a named entity recognition (NER) model’s output is input to a relation classification (RC) model; (2) joint sequence-to-sequence style models where the raw input text is directly transformed into relations through linearization schemas; and (3) generative large language models (LLMs), where prompts, fine-tuning, and in-context learning are being leveraged for RE. While LLMs are becoming popular because of tools such as ChatGPT, the biomedical NLP community needs to carefully evaluate which paradigm is more suitable for E2ERE. In this effort, using the RareDis dataset as a complex use-case, we evaluate the best representative models from each of the three paradigms for E2ERE. Our findings reveal that pipeline models are still the best, while sequence-to-sequence models are not far behind. We verify these findings on a second E2ERE dataset for chemical-protein interactions. Although LLMs are more suitable for zero-shot settings, our results show that it is better to work with more conventional models trained and tailored for E2ERE when training data is available. Our contribution is also the first to conduct E2ERE for the RareDis dataset.
KW - biomedical relation extraction
KW - encoder-decoder models
KW - end-to-end models
KW - information extraction
KW - large language models
UR - https://www.scopus.com/pages/publications/105010826954
UR - https://www.scopus.com/pages/publications/105010826954#tab=citedBy
U2 - 10.1007/978-3-031-97141-9_4
DO - 10.1007/978-3-031-97141-9_4
M3 - Conference contribution
AN - SCOPUS:105010826954
SN - 9783031971402
VL - 15836
T3 - Lecture Notes in Computer Science
SP - 49
EP - 63
BT - Natural Language Processing and Information Systems - 30th International Conference on Applications of Natural Language to Information Systems, NLDB 2025, Proceedings
A2 - Ichise, Ryutaro
T2 - 30th International Conference on Natural Language and Information Systems, NLDB 2025
Y2 - 4 July 2025 through 6 July 2025
ER -