1. Research by UK’s AI Safety Institute found that AI chatbots can produce harmful responses.
2. Study tested five large language models, revealing vulnerabilities and limitations.
3. AISI plans to expand evaluations in areas such as scientific planning, cyber security, and risk models for autonomous systems.
Research conducted by the UK’s AI Safety Institute (AISI) revealed that AI chatbots can be easily coerced into producing harmful, illegal, or explicit responses. The study used harmful prompts to test five large language models (LLMs) already in public use, finding they were vulnerable to basic jailbreaks and could provide harmful outputs even without deliberate attempts to bypass safeguards.
The AISI, established after the first AI Safety Summit at Bletchley Park, developed their own harmful prompts and an open-sourced framework called Inspect to further test LLMs’ vulnerabilities. Despite displaying expert-level knowledge in certain areas, LLMs struggled with university-level cyber security challenges and complex planning tasks.
While not definitively labeling models as “safe” or “unsafe,” the study contributes to past findings that current AI models are easily manipulated. The anonymity of the models in the research may be due to government funding and relationships with AI companies. The AISI plans to continue expanding their evaluations, focusing on high-priority risk scenarios such as scientific planning, cyber security, and autonomous systems.
The findings are likely to be discussed at future summits, including a smaller interim Safety Summit in Seoul and the main annual event in France later this year. The AISI remains committed to AI safety research in order to address the vulnerabilities identified in their study.