Ir directamente a la navegación principal Ir directamente a la búsqueda Ir directamente al contenido principal

Comparison of Pipelines, Seq2seq Models, and LLMs for Rare Disease Information Extraction

Producción científica: Conference contributionrevisión exhaustiva

Resumen

End-to-end relation extraction (E2ERE) is an important application of natural language processing (NLP) in biomedicine. The extracted relations populate knowledge graphs and drive more high level applications in knowledge discovery and information retrieval. E2ERE is frequently handled at the sentence level involving continuous entities. A more complex setting is document level E2ERE with discontinuous and overlapping/nested entities. We identified a recently introduced RE dataset for rare diseases (RareDis) that has these complex traits. Among current E2ERE methods, we see three well-known paradigms: (1) pipeline based approaches where a named entity recognition (NER) model’s output is input to a relation classification (RC) model; (2) joint sequence-to-sequence style models where the raw input text is directly transformed into relations through linearization schemas; and (3) generative large language models (LLMs), where prompts, fine-tuning, and in-context learning are being leveraged for RE. While LLMs are becoming popular because of tools such as ChatGPT, the biomedical NLP community needs to carefully evaluate which paradigm is more suitable for E2ERE. In this effort, using the RareDis dataset as a complex use-case, we evaluate the best representative models from each of the three paradigms for E2ERE. Our findings reveal that pipeline models are still the best, while sequence-to-sequence models are not far behind. We verify these findings on a second E2ERE dataset for chemical-protein interactions. Although LLMs are more suitable for zero-shot settings, our results show that it is better to work with more conventional models trained and tailored for E2ERE when training data is available. Our contribution is also the first to conduct E2ERE for the RareDis dataset.

Idioma originalEnglish
Título de la publicación alojadaNatural Language Processing and Information Systems - 30th International Conference on Applications of Natural Language to Information Systems, NLDB 2025, Proceedings
EditoresRyutaro Ichise
Páginas49-63
Número de páginas15
DOI
EstadoPublished - 2026
Evento30th International Conference on Natural Language and Information Systems, NLDB 2025 - Kanazawa, Japan
Duración: jul 4 2025jul 6 2025

Serie de la publicación

NombreLecture Notes in Computer Science
Volumen15836 LNCS
ISSN (versión impresa)0302-9743
ISSN (versión digital)1611-3349

Conference

Conference30th International Conference on Natural Language and Information Systems, NLDB 2025
País/TerritorioJapan
CiudadKanazawa
Período7/4/257/6/25

Nota bibliográfica

Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.

ASJC Scopus subject areas

  • Theoretical Computer Science
  • General Computer Science

Huella

Profundice en los temas de investigación de 'Comparison of Pipelines, Seq2seq Models, and LLMs for Rare Disease Information Extraction'. En conjunto forman una huella única.

Citar esto