Abstract
Trusted large language models (LLMs) inherit ethical guidelines to prevent generating harmful content, whereas malicious LLMs are engineered to enable the generation of unethical and toxic responses. Both trusted and malicious LLMs use guardrails in differential contexts per the requirements of the developers and attackers, respectively. We explore the multifaceted world of guardrails implementation in LLMs by conducting an empirical analysis to assess the effectiveness of guardrails using prompts. Our results revealed that guardrails deployed in the trusted LLMs could be bypassed using prompt manipulation techniques such as “pretend” and “persist” to generate harmful content. In addition, we also discovered that malicious LLMs still deploy weak guardrails to evade detection by generating human-like content. This empirical analysis provides insights into the design of the malicious and trusted LLMs. We also propose recommendations to defend against prompt manipulation and guardrails bypass while designing LLMs.
| Original language | English |
|---|---|
| Pages (from-to) | 98-108 |
| Number of pages | 11 |
| Journal | IEEE Technology and Society Magazine |
| Volume | 44 |
| Issue number | 3 |
| DOIs | |
| State | Published - 2025 |
Bibliographical note
Publisher Copyright:© 1982-2012 IEEE.
Funding
The work of Sherali Zeadally was supported in part by the Distinguished Visiting Professorship from the University of Johannesburg, South Africa.
| Funders |
|---|
| University of Johannesburg |
ASJC Scopus subject areas
- General Engineering
- General Social Sciences
Fingerprint
Dive into the research topics of 'The Unprecedented Surge in Generative AI: Empirical Analysis of Trusted and Malicious Large Language Models (LLMs)'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver