LIVE Watch Now
Breaking
Ancient Inca Mummies Reveal Smallpox Originated from EuropeFederal Minister Chairs Meeting with Saudi InvestorsECP rejects Achakzai as PkMAP leader over election failureBrother of Uzair Baloch Shot in Karachi’s LyariPTA Urges Parents to Monitor Children’s Social Media UseGovernment Hikes Petrol and Diesel Prices AgainKhyber Pakhtunkhwa Launches AI Training Program for Teachers and OfficersRussian missile strikes US-owned drone factory in KyivPM Praises Security Forces for Khuzdar Operation SuccessPM Shehbaz Orders Technical Audit of DISCOs’ Billing SystemPakistan, Iran Pledge Enhanced Cooperation in Multiple SectorsGoogle DeepMind’s Gemini Robotics 2 Enables Whole-Body Control for Humanoid RobotsSindh Schools Return to Six-Day Week Starting August 1PM Shehbaz Urges Digitisation of Consular Services for Overseas PakistanisPSX Plummets as US-Iran Tensions Drag Down MarketPakistani Celebrities Unite for Peace Amid AJK TensionsEU to Fund Seven AI Gigafactories with €10 Billion PlanChina denies reports of missile supply to IranSSUET Job Fair Sees Over 120 Companies Participate as AI Lab InauguratedVocational training in CPEC 2.0 program transforms Pakistani studentsAncient Inca Mummies Reveal Smallpox Originated from EuropeFederal Minister Chairs Meeting with Saudi InvestorsECP rejects Achakzai as PkMAP leader over election failureBrother of Uzair Baloch Shot in Karachi’s LyariPTA Urges Parents to Monitor Children’s Social Media UseGovernment Hikes Petrol and Diesel Prices AgainKhyber Pakhtunkhwa Launches AI Training Program for Teachers and OfficersRussian missile strikes US-owned drone factory in KyivPM Praises Security Forces for Khuzdar Operation SuccessPM Shehbaz Orders Technical Audit of DISCOs’ Billing SystemPakistan, Iran Pledge Enhanced Cooperation in Multiple SectorsGoogle DeepMind’s Gemini Robotics 2 Enables Whole-Body Control for Humanoid RobotsSindh Schools Return to Six-Day Week Starting August 1PM Shehbaz Urges Digitisation of Consular Services for Overseas PakistanisPSX Plummets as US-Iran Tensions Drag Down MarketPakistani Celebrities Unite for Peace Amid AJK TensionsEU to Fund Seven AI Gigafactories with €10 Billion PlanChina denies reports of missile supply to IranSSUET Job Fair Sees Over 120 Companies Participate as AI Lab InauguratedVocational training in CPEC 2.0 program transforms Pakistani students
◕ SundialUpdated 16 hours ago
Trending Stories
Science & Health

Researchers Warn of Inherent Security Flaws in LLMs

Researchers warn that large language models may be inherently insecure against prompt injection attacks due to their inability to distinguish between trust

Add Sun News on Google News
Researchers Warn of Inherent Security Flaws in LLMs
Photo: ProPakistani

Key Takeaways

  • Large language models (LLMs) may be fundamentally insecure against prompt injection attacks.
  • Studies show that LLMs struggle to distinguish between trusted instructions and their own reasoning.
  • Attackers can exploit this weakness by mimicking internal reasoning, bypassing model safeguards.

Researchers have warned of a fundamental security flaw in large language models (LLMs) that could make them inherently vulnerable to prompt injection attacks. This issue was highlighted at the International Conference on Machine Learning (ICML), where a paper titled 'Prompt Injection as Role Confusion' presented findings indicating that LLMs are unable to reliably distinguish between trusted instructions, user requests, external data, and their own reasoning.

The study found that models often judge roles based more on the wording and style of text than on tags surrounding it. This means that text written to look like internal reasoning could be treated as trusted even when it originates from a user. For instance, changing the tags around a passage made little difference if the writing still resembled a particular role.

One specific attack technique called 'chain-of-thought forgery' was demonstrated by researchers. In this method, attackers present fabricated reasoning written in a style that mimics the model's internal notes. The model then interprets this fake reasoning as something it has already concluded and acts upon it. This approach allowed researchers to bypass safeguards on several models, making them provide information they had been trained not to give.

AdvertisementFollow Sun NewsWatch Live

In one test, irrelevant details about the user’s clothing were combined with fabricated internal reasoning that claimed these details made an otherwise prohibited request acceptable. OpenAI models tested by the researchers then complied with the request, demonstrating a success rate of approximately 60% across multiple models.

The attack technique was so effective that it won an OpenAI red-teaming competition in August 2025. This event highlighted the severity of the issue and underscored the need for better security measures. However, one author of the paper, Jasmine Cui, argues that current safety training methods may be a game of 'whack-a-mole,' where developers can teach models many examples of what they should not do but cannot anticipate every possible attack.

The implications of this research are significant as LLMs are increasingly used in government systems, military applications, healthcare, online shopping, and other critical areas. Incorrect or manipulated actions by these models could have serious consequences. The findings suggest that better safety training alone may never fully eliminate the problem, as it is built into how current models process text.

AI companies regularly use human red teams and automated systems to find new attacks before models are released. Once researchers discover a successful jailbreak or prompt injection, developers can train later models to recognize and resist similar attacks. However, this approach has its limitations, as the paper's authors argue that it is difficult to anticipate every possible attack scenario.

The research highlights the ongoing challenges in ensuring the security of AI systems and underscores the need for continued innovation and vigilance in the field.

Sources