Former Google researchers plan to build an AI "referee" system to prevent the risk of models getting out of control

📅 2026-08-26

Abstract:

Be at the heart of the development of artificial intelligence to ensure this powerful technology remains safe and reduce the risk of it slipping out of the control of its creators. This is partly a response to common practices in the tech industry today. In many cases, businesses are leveraging AI to regulate AI. Automated approaches have efficiency advantages, and research shows that AI may be better than humans at finding model vulnerabilities. As AI capabilities continue to increase, this gap is likely to widen further.

But Rishub Jain, who co-founded Sampura Research, believes that humans still need to play an important role.

He and his co-founders have raised US$6.5 million in funding and received an additional US$4.2 million in committed investments. They plan to develop a hybrid human-machine system called "Judge" to help companies identify and block security vulnerabilities that AI models try to exploit, thereby reducing the risk of them circumventing supervision or even losing control. The work has become more urgent since OpenAI and Anthropic disclosed that their models had hacked into other companies' systems without authorization.

“I think we still have a lot to learn from humans,” Jain said. "If humans are involved in the entire process from the beginning, the chances of the model learning to circumvent these vulnerabilities will be less."

As Alphabet's Google, Anthropic and OpenAI fiercely compete for dominance in the artificial intelligence era, some researchers within the lab believe that commercial pressure is outweighing security considerations. Many researchers have resigned in protest as a result, warning that AI is developing too rapidly or is being used for purposes they cannot accept.

Related tags

Related articles

Comments

0/500
验证码
No comments yet