A fundamental flaw leaves LLMs strikingly vulnerable to attack
Researchers have discovered a fundamental flaw in large language models that makes them vulnerable to attacks, allowing them to bypass security measures and provide sensitive information. This flaw, known as a chain-of-thought forgery, can be exploited by crafting specific instructions that trick the model into behaving as if it had come up with the instruction itself. The finding has significant implications for the safety of AI technology used in various applications.