Resumen
Trusted large language models (LLMs) inherit ethical guidelines to prevent generating harmful content, whereas malicious LLMs are engineered to enable the generation of unethical and toxic responses. Both trusted and malicious LLMs use guardrails in differential contexts per the requirements of the developers and attackers, respectively. We explore the multifaceted world of guardrails implementation in LLMs by conducting an empirical analysis to assess the effectiveness of guardrails using prompts. Our results revealed that guardrails deployed in the trusted LLMs could be bypassed using prompt manipulation techniques such as “pretend” and “persist” to generate harmful content. In addition, we also discovered that malicious LLMs still deploy weak guardrails to evade detection by generating human-like content. This empirical analysis provides insights into the design of the malicious and trusted LLMs. We also propose recommendations to defend against prompt manipulation and guardrails bypass while designing LLMs.
| Idioma original | English |
|---|---|
| Páginas (desde-hasta) | 98-108 |
| Número de páginas | 11 |
| Publicación | IEEE Technology and Society Magazine |
| Volumen | 44 |
| N.º | 3 |
| DOI | |
| Estado | Published - 2025 |
Nota bibliográfica
Publisher Copyright:© 1982-2012 IEEE.
Financiación
The work of Sherali Zeadally was supported in part by the Distinguished Visiting Professorship from the University of Johannesburg, South Africa.
| Financiadores |
|---|
| University of Johannesburg |
ASJC Scopus subject areas
- General Engineering
- General Social Sciences
Huella
Profundice en los temas de investigación de 'The Unprecedented Surge in Generative AI: Empirical Analysis of Trusted and Malicious Large Language Models (LLMs)'. En conjunto forman una huella única.Citar esto
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver