Abstract:
OpenAI plans to allow third-party organizations to assess security risks earlier in the development cycle of artificial intelligence (AI) models, a move aimed at continuing efforts to address concerns surrounding the potential harm of AI. The ChatGPT developer will announce in a blog post on Tuesday that it plans to have external agencies conduct technical security assessments during the training, evaluation and rollout of new AI models.

OpenAI also lists the priorities it believes are needed for such assessments to be effective, including "strong independent mechanisms, scientific rigor, robust security practices and clear accountability."
Lama Ahmad, who is responsible for most of the company's security review work in cooperation with external experts, said that previously, OpenAI usually mainly introduced such institutions to conduct security assessments and evaluate model capabilities before model release.
“As the potential risks and impacts rise, in addition to deployment, we also want to ensure that critical aspects such as training and evaluation are looked at,” Ahmad said in an interview.
For some of the most sensitive jobs, OpenAI may let outside evaluators come into the company's premises to conduct assessments, a practice the company has tried in the past, Ahmad said.
Dario Amodei, CEO of OpenAI competitor Anthropic PBC, recently called on the AI industry to support slowing down development and said that new security measures will be adopted, including the introduction of third-party evaluation agencies. OpenAI CEO Sam Altman later endorsed the approach. Anthropic said late last week that it would bring in Accenture evaluators to test the safety of its cutting-edge AI models.
OpenAI said in a blog post that the company is in discussions with institutions that may be involved in such assessments. Ahmad said that the negotiation targets include both institutions that have cooperated in the past and institutions that have not cooperated before, including AI research institutions METR and Redwood Research. OpenAI had previously commissioned these two institutions to investigate the intrusion of its model into the Hugging Face system.
"AI Labs have a responsibility to create conditions for effective review while protecting sensitive information," OpenAI said in a blog post.
Comments