Ir directamente a la navegación principal Ir directamente a la búsqueda Ir directamente al contenido principal

CrunchLLM: Multitask LLMs for structured business reasoning and outcome prediction

Producción científica: Articlerevisión exhaustiva

Resumen

Predicting the success of startup companies, defined as achieving an exit through acquisition or IPO, is a critical problem in entrepreneurship and innovation research. Datasets such as Crunchbase provide both structured information (e.g., funding rounds, industries, and investor networks) and unstructured text (e.g., company descriptions), but effectively leveraging such heterogeneous data for prediction remains challenging. Traditional machine learning approaches often rely only on structured features and achieve moderate accuracy, while large language models (LLMs) offer strong reasoning capabilities but are not readily adapted to domain-specific business data. We present CrunchLLM, a domain-adapted and backbone-agnostic LLM framework for startup success prediction. CrunchLLM integrates structured company attributes with unstructured textual narratives and applies parameter-efficient fine-tuning together with prompt optimization to specialize foundation models for entrepreneurship data. Importantly, our framework introduces a self-verifiable multitask objective, in which the justification loss serves as a training-time constraint on classification, together with a hierarchically ordered input encoding that reduces the tendency of long unstructured company narratives to overshadow structured business attributes. These methodological innovations yield more reliable and feature-grounded predictions than conventional prompt-based LLM adaptation. Our approach achieves 89% accuracy on the Crunchbase startup success prediction task, significantly outperforming traditional classifiers and baseline LLMs. Beyond predictive performance, CrunchLLM generates interpretable reasoning traces that support its predictions, enhancing transparency and trustworthiness for financial and policy decision-makers. Overall, this work demonstrates how domain-aware LLM adaptation and structured–unstructured data fusion can advance predictive modeling of entrepreneurial outcomes, providing both a methodological framework and a practical tool for data-driven decision-making in venture capital and innovation policy.

Idioma originalEnglish
Número de artículo133754
PublicaciónNeurocomputing
Volumen686
DOI
EstadoPublished - jul 14 2026

Nota bibliográfica

Publisher Copyright:
© 2026 Elsevier B.V.

Financiación

The authors wish to thank Andrew Cheng (Carnegie Mellon University) and Qiang Ye (University of Kentucky) for their insightful discussions and intellectual contributions regarding the theoretical framework of self-verifiable learning and the formal proof of Theorem 1 provided during the revision stage. This research is supported in part by the NSF under Grant IIS 2327113 and ITE 2433190 , the NIH under Grants R21AG070909 and P30AG072946 , the National Artificial Intelligence Research Resource (NAIRR) Pilot NSF OAC 240219, and Jetstream2, Bridges2, and Neocortex resources. We thank the University of Kentucky Center for Computational Sciences and Information Technology Services Research Computing for their support and use of the Lipscomb Compute Cluster and associated research computing resources. Qiang Cheng received the BS and MS degrees from Peking University, China, and the Ph.D. degree from the Department of Electrical and Computer Engineering at the University of Illinois, Urbana-Champaign. He is an associate professor at the Institute for Biomedical Informatics and the Department of Computer Science at the University of Kentucky. He previously worked as a faculty fellow at the Air Force Research Laboratory, Wright-Patterson, OH, and a senior researcher at Siemens Corporate Research and Siemens Medical Solutions, Siemens Corp., Princeton, NJ. His research interests include data science, machine learning, pattern recognition, and biomedical informatics. Supported by NSF, ARO, and NIH, he has published around 200 peer-reviewed papers in various premium venues, including IEEE TPAMI, TNNLS, TSP, NIPS, CVPR, AAAI, ACM TIST, TKDD, and KDD. He has a number of international patents issued or filed with the University of Kentucky, IBM T.J. Watson Research Laboratory, and Siemens Medical.

FinanciadoresNúmero del financiador
Mellon College of Science, Carnegie Mellon University
ARO MURI
Kentucky Transportation Center, University of Kentucky
National Institutes of Health (NIH)P30AG072946, R21AG070909
National Artificial Intelligence Research ResourceBridges2, Jetstream2, OAC 240219
National Science Foundation Arctic Social Science ProgramITE 2433190, IIS 2327113

    ASJC Scopus subject areas

    • Computer Science Applications
    • Cognitive Neuroscience
    • Artificial Intelligence

    Huella

    Profundice en los temas de investigación de 'CrunchLLM: Multitask LLMs for structured business reasoning and outcome prediction'. En conjunto forman una huella única.

    Citar esto