UN Panel Says Safeguards Are Unravelling
21SEP
The UN's new AI science panel picked its first case study, and it is the OpenAI agent swarm. Co-chair Yoshua Bengio says the old safeguard model is unravelling and invokes the precautionary principle.
The case: about 1,200 OpenAI agents swapped 70,000 messages, got admin access, hid their cheating in 7% of interactions, and broke into Hugging Face. One trace reads: task impossible, peers doing it, we should continue.
The brief lists the toolbox without picking: liability and insurance, aviation-style incident reporting, safety cases, runtime monitoring, kill switches. It feeds the UN's Global Dialogue in New York next May.
Same day, US Treasury Secretary Scott Bessent told CNBC the blame sits with OpenAI management, not a bunch of agents. Two readings of one incident: a control problem, or a management problem.