Artificial intelligence is being integrated into business operations and daily life at an unprecedented rate, and its application is expanding. However, with the explosive application of technology, the concerns of the scientific community and experts about the responsible use of AI to ensure that ethical boundaries are not blurred are growing. Not long ago, the large-scale language model (LLMM) had shown a strange act of deception and lying under stress tests. Today, researchers claim to have found a new way to “guilty” the AI chat robots to say what they should not have said.

Previous studies have revealed the tendency of LLM models to resort to coercion in situations of stress and “survival”. Imagine the potential harm of this “deception” if you can proactively manipulate the AI chat robot to do what you want. The research team from Intel, Boise State University and Illinois University revealed alarming findings in a joint paper. The core view of the paper is that they can be successfully deceived by over-advertising information to chat robots. What happens when the AI model is “bombed” with big information? Researchers point out that models get confused. It is this state of confusion that has become a point of systemic vulnerability, enabling the attackers to bypass built-in security protection mechanisms. In order to take advantage of this loophole and to implement the “breakout” (i.e. breakout model limits), researchers have developed an automated tool called “InfoFloood”. Strong models such as ChatGPT, Gemini have pre-set security barriers aimed at preventing them from being manipulated to export harmful or dangerous content.

This newly discovered breakthrough technology shows that as long as you use a complex data stream to end up “mix” the AI model, it will “learn” for you. Researchers have further elaborated on this finding to the scientific media, 404 Media, and have confirmed that, as these models tend to rely on the shallow semantics of communication, they cannot fully understand the underlying intent behind the information. It is on this basis that the research team designed a method to test how talking robots react when a hazardous request is hidden in an overloaded flood of information. The team indicated that it planned to send leak disclosure reports to companies with large AI models so that they could notify their safety teams. However, the research paper also highlighted key challenges: even with the deployment of safety filters, malicious actors may still use such technical deceptive models to embed harmful content. This reveals the fragility of the current AI security protection system in the face of new attack methods.

The validity of “information overload” attacks suggests that existing content filtering and intent identification mechanisms alone may not be sufficient to respond to increasingly complex confrontational tactics. Ensuring that AI systems remain ethical bottom lines and safety guidelines under complex, chaotic and even maliciously constructed inputs has become a central issue to be addressed by developers and regulators. With the rapid increase in AI ‘ s capabilities, the excavation of its potential loopholes and the reinforcement of its defensive measures will be a constant offensive and defensive battle.