Researchers Develop Method to Unalign Safety-Aligned AI Models with Single Unlabeled Prompt
A team of researchers has introduced GRP-Obliteration, a method that can unalign safety-aligned AI models using a single unlabeled prompt, potentially undermining their robustness. The technique, which leverages Group Relative Policy Optimization, can reliably remove safety constraints from target models while preserving their utility. Lire la suite


















