This article was written with the assistance of AI.
News Factory APP - agentic news to boost your SEO & AEO.
Open-Weight LLMs Lead in Resisting Russian Propaganda, Study Finds
Key Points
- Open-weight LLMs like Nvidia's Nemotron and Alibaba's Qwen outperformed many proprietary models in resisting Russian propaganda.
- OpenAI's GPT-5.4 achieved the highest mean score of 88.9, with 54% of responses marked "Exemplary."
- Claude 3.5 Haiku (2024) scored 73.1, placing it in the bottom third of 2026 models on the benchmark.
- Google's Gemini 2.5 Pro scored 82 but showed vulnerability to malicious prompts; Gemini 3.5 Flash fell to 73.
- Performance drops sharply when models are tested in Russian, affecting Gemini 3.5 Flash, Moonshot's Kimi K2, and StepFun's Step 3.5 Flash.
- Research by Gregory Asmolov warns that Russia is attempting to steer AI outputs through BRICS collaborations.
- The benchmark underscores the need for stronger multilingual defenses against state‑sponsored disinformation.