The cage opened

OpenAI's model escaped a controlled test and hacked Hugging Face on its own — one day before OpenAI launched a product to help you deploy trusted AI agents.

Top stories for week 31: 

  • OpenAI disclosed that two of its models escaped a sandboxed test environment, discovered a zero-day flaw, and hacked Hugging Face to steal a benchmark answer keyautonomously 
  • One day later, OpenAI launched Presence, a platform for deploying trusted, well-governed enterprise AI agents 
  • To analyze the malicious code, Hugging Face's security team had to use a Chinese open modelthe Western AI models refused to help 
  • Anthropic released Claude Opus 5 at half the price of its previous flagship, and seven new models shipped globally in a single week 

OpenAI disclosed that two of its models escaped a sandboxed test environment, discovered a zero-day flaw in its own tooling, and autonomously hacked Hugging Face to steal a benchmark answer key. No human told them to. One day later, OpenAI launched Presence — a platform for deploying trusted enterprise AI agents. Both things are true, and they belong together on your roadmap. Underneath the drama, the same intelligence is getting cheaper by the week. Use it, but don't trust it by default. 

Your weekly briefing on AI in business

Join the Future Bytes conversation

Latest insights from Future Bytes