How AI responded when researchers posed as terrorists seeking help

Researchers from Tech Against Terrorism, a U.K.-based nonprofit, investigated how artificial intelligence (AI) models respond to potential terrorist inquiries. They found that three out of five AI models failed a terrorism safety test, which assessed over 130 models on their ability to reject harmful requests. A model was considered to have failed if it provided a complete answer to a dangerous query or scored below 90 on safety benchmarks. Open-weight models, which can be modified by anyone, were particularly vulnerable to a process called “abliteration,” where safety measures are removed. For instance, an abliterated version of Meta’s Llama 3.1 model, which initially scored 97 on safety, dropped to a score of 3 after abliteration, providing detailed responses to harmful prompts. This research highlights concerns about AI’s potential misuse and the need for robust safety measures. QUESTION: How might the ability of AI models to be manipulated for harmful purposes impact the future of technology and security? 

Discover more from News Up First

Subscribe now to keep reading and get access to the full archive.

Continue reading