Skip to main navigation Skip to search Skip to main content

The Unprecedented Surge in Generative AI: Empirical Analysis of Trusted and Malicious Large Language Models (LLMs)

Research output: Contribution to journalArticlepeer-review

Abstract

Trusted large language models (LLMs) inherit ethical guidelines to prevent generating harmful content, whereas malicious LLMs are engineered to enable the generation of unethical and toxic responses. Both trusted and malicious LLMs use guardrails in differential contexts per the requirements of the developers and attackers, respectively. We explore the multifaceted world of guardrails implementation in LLMs by conducting an empirical analysis to assess the effectiveness of guardrails using prompts. Our results revealed that guardrails deployed in the trusted LLMs could be bypassed using prompt manipulation techniques such as “pretend” and “persist” to generate harmful content. In addition, we also discovered that malicious LLMs still deploy weak guardrails to evade detection by generating human-like content. This empirical analysis provides insights into the design of the malicious and trusted LLMs. We also propose recommendations to defend against prompt manipulation and guardrails bypass while designing LLMs.

Original languageEnglish
Pages (from-to)98-108
Number of pages11
JournalIEEE Technology and Society Magazine
Volume44
Issue number3
DOIs
StatePublished - 2025

Bibliographical note

Publisher Copyright:
© 1982-2012 IEEE.

Funding

The work of Sherali Zeadally was supported in part by the Distinguished Visiting Professorship from the University of Johannesburg, South Africa.

Funders
University of Johannesburg

    ASJC Scopus subject areas

    • General Engineering
    • General Social Sciences

    Fingerprint

    Dive into the research topics of 'The Unprecedented Surge in Generative AI: Empirical Analysis of Trusted and Malicious Large Language Models (LLMs)'. Together they form a unique fingerprint.

    Cite this