Ir directamente a la navegación principal Ir directamente a la búsqueda Ir directamente al contenido principal

The Unprecedented Surge in Generative AI: Empirical Analysis of Trusted and Malicious Large Language Models (LLMs)

Producción científica: Articlerevisión exhaustiva

Resumen

Trusted large language models (LLMs) inherit ethical guidelines to prevent generating harmful content, whereas malicious LLMs are engineered to enable the generation of unethical and toxic responses. Both trusted and malicious LLMs use guardrails in differential contexts per the requirements of the developers and attackers, respectively. We explore the multifaceted world of guardrails implementation in LLMs by conducting an empirical analysis to assess the effectiveness of guardrails using prompts. Our results revealed that guardrails deployed in the trusted LLMs could be bypassed using prompt manipulation techniques such as “pretend” and “persist” to generate harmful content. In addition, we also discovered that malicious LLMs still deploy weak guardrails to evade detection by generating human-like content. This empirical analysis provides insights into the design of the malicious and trusted LLMs. We also propose recommendations to defend against prompt manipulation and guardrails bypass while designing LLMs.

Idioma originalEnglish
Páginas (desde-hasta)98-108
Número de páginas11
PublicaciónIEEE Technology and Society Magazine
Volumen44
N.º3
DOI
EstadoPublished - 2025

Nota bibliográfica

Publisher Copyright:
© 1982-2012 IEEE.

Financiación

The work of Sherali Zeadally was supported in part by the Distinguished Visiting Professorship from the University of Johannesburg, South Africa.

Financiadores
University of Johannesburg

    ASJC Scopus subject areas

    • General Engineering
    • General Social Sciences

    Huella

    Profundice en los temas de investigación de 'The Unprecedented Surge in Generative AI: Empirical Analysis of Trusted and Malicious Large Language Models (LLMs)'. En conjunto forman una huella única.

    Citar esto