Jailbreaking the powerful GenAI models
Being the most powerful Generative AI doesn’t generally mean that they are safe in protecting sensitive data. Their guardrails are not foolproof, cybersecurity experts say. ChatGPT 5, the most powerful LLM upgrade from OpenAI, seems to be no exception. This creates a significant danger for organisations where these tools are being rapidly adopted by employees, often without oversight.
Just 24 hours after OpenAI launched its highly anticipated GPT-5 model, which is supposed to offer sophisticated prompt safety, exposure management company Tenable has claimed that it successfully jailbroke the platform. It stated that the GenAI engine provided detailed instructions on how to build a Molotov cocktail, a handheld weapon studded with inflammable substances.
Tenable said that its researchers used a social engineering method known as the crescendo technique to bypass these safety protocols in just four simple prompts by posing as a history student interested in the historical context and recipe of the incendiary device.
“The successful jailbreak highlights a critical security gap in the latest generation of AI models, demonstrating that despite developer claims, they remain vulnerable to manipulation for malicious purposes,” Tomer Avni, Vice-President (Product Management) at Tenable, said.
“The ease with which we bypassed GPT-5’s new safety protocols proves that even the most advanced AI is not foolproof,” he said.
Without proper visibility and governance, businesses are unknowingly exposed to serious security, ethical, and compliance risks. This incident is a clear call for a dedicated AI exposure management strategy to secure every model in use.
These jailbreaks prove that organisations cannot rely solely on the built-in safety features offered by the LLMs.
Jailbreaking DeepSeek
A similar jailbreak was reported by the Unit 42 research team of Palo Alto Networks. The research team jailbroke the open-source reasoning model DeepSeek recently.
Swapna Bapat, Vice-President and Managing Director (India and SAARC) of Palo Alto Networks, said that the team used multiple techniques to bypass its built-in safeguards. While initial responses often appeared benign, carefully crafted follow-up prompts elicited detailed malicious outputs, ranging from phishing kits to attack instructions for techniques such as SQL injection and lateral movement.
Unit 42 researchers used two tested jailbreaking techniques (Deceptive Delight and Bad Likert Judge) and a new method called Crescendo against DeepSeek models. “We achieved significant bypass rates, with little to no specialised knowledge or expertise being necessary,” the team said.
“Our research findings show that these jailbreak methods can elicit explicit guidance for malicious activities. These activities include data exfiltration tooling, keylogger creation and even instructions for incendiary devices, demonstrating the tangible security risks posed by this emerging class of attack,” it said.
While it can be challenging to guarantee complete protection against all jailbreaking techniques for a specific LLM, organisations can implement security measures that can help monitor when and how employees are using LLMs. This becomes crucial when employees are using unauthorised third-party LLMs, a cybersecurity expert said.
Published on August 14, 2025

Leave a Comment