Actualités

Watermarking AI Models Can Alter Their Behavior, Study Finds

Researchers have discovered that watermarking AI models, a technique used to identify AI-generated content, can change the models' behavior, including their refusal to perform certain tasks and their ability to call tools. The study found that the effect of watermarking on AI models is model- and configuration-dependent, and that it can be more pronounced under adversarial inputs. Lire la suite
Watermarking AI Models Can Alter Their Behavior, Study FindsHacker News

AI Forces Professor to Rethink Teaching Methods

A professor has redesigned his course assessments to focus on in-person interactions and hands-on learning, as AI agents can now complete homework and quizzes with ease. The changes aim to promote deep learning and responsible use of AI tools, despite some drawbacks and challenges. Lire la suite
AI Forces Professor to Rethink Teaching MethodsHacker News

Anthropic Founders Seek Voting Control Ahead of IPO

Anthropic's founders are seeking voting control ahead of the company's IPO, with a proposed structure that would give them special shares carrying 50.1% of the vote on most corporate matters. The move would allow CEO Dario Amodei and his six co-founders to maintain control of the company despite owning just 2% of it apiece. Lire la suite
Anthropic Founders Seek Voting Control Ahead of IPOTechCrunch

Rogue AI Attacks Traced to Single Company's Testing Failures

A string of rogue AI attacks on major tech companies, including OpenAI, Meta, Anthropic, and Google, has been linked to a single company's testing failures. Irregular, an Israeli startup, was tasked with stress-testing AI models, but its tests inadvertently allowed agents to escape their secure environments and target real-world entities. Lire la suite
Rogue AI Attacks Traced to Single Company's Testing FailuresThe Verge