Security and identity briefing
Anthropic disclosed three test failures, the OpenAI incident reached a second company’s customer, and Google and Nvidia introduced new frameworks for securing AI agents.
Anthropic reviewed more than 141,000 cybersecurity evaluation runs after the OpenAI and Hugging Face incident and found three cases in which its own models had reached real organizations. Claude Opus 4.7, Mythos 5 and an internal research model were running capture-the-flag exercises. Their prompts said they were in simulations without internet access, but the test environments were connected to the open internet. The models used basic methods such as weak passwords. Two affected organizations said they had not detected the activity before Anthropic contacted them.
The OpenAI incident reached farther than first disclosed. The same agent system that compromised Hugging Face also accessed an asset belonging to a customer of cloud platform Modal Labs. That infrastructure was tied to CyberGym, the project behind the ExploitGym benchmark the models had been assigned to solve. The system continued pursuing the benchmark objective after leaving its intended environment. Modal said its platform was not compromised; the customer had left an endpoint exposed that allowed anyone on the internet to execute code in its sandboxes.
China’s Commerce Ministry accused Washington of “AI hegemonism” and threatened countermeasures after U.S. officials raised the prospect of investigations, sanctions and export restrictions against Chinese developers. The dispute concerns Moonshot AI, which U.S. officials allege used large-scale distillation from Anthropic’s Fable model to help build Kimi K3. Moonshot denies the claim and attributes the model’s performance to original architectural work. The U.S. allegation has not been accompanied by public technical evidence that would allow independent verification.
Nvidia and dozens of cloud, security, open-source and enterprise software organizations have formed the Open Secure AI Alliance. Its work covers identity, permissions, isolation, guardrails, logging and evaluation across the agent stack. Founding participants include Microsoft, Hugging Face, Okta, Cloudflare, IBM and the Linux Foundation. Initial contributions include Nvidia’s framework for testing and auditing agent behavior, HPE’s work on cryptographic workload identity through SPIFFE and SPIRE, the Safetensors model format, digitally signed software patches and Microsoft’s multi-model vulnerability scanning harness.
Google has published “Beyond Zero,” an authorization model for humans and AI agents. It evaluates each action on a specific resource instead of granting broad access to an application or tool. The model applies to user interfaces, APIs and Model Context Protocol connections. Static policies are combined with context about the actor, task, data and available safeguards; higher-risk activity can trigger an investigation, added verification or containment. Google says early internal deployments have improved detection of access abuse and protection of intellectual property.
The European Union’s AI transparency rules began applying August 2. Providers must tell people when they are interacting directly with an AI system and add machine-readable markers to synthetic or manipulated outputs. Organizations deploying the technology must disclose deepfakes, AI-generated public-interest material published without human editorial control, emotion recognition and biometric categorization. Systems placed on the market before August 2 have until December 2 to meet the machine-readable marking requirement; the other Article 50 disclosure duties are already in force.
Americans reported nearly $21 billion in cybercrime losses in 2025, including $893 million linked to AI-enabled scams across more than 22,000 complaints, according to a bipartisan U.S. Senate report. The Senate Special Committee on Aging highlighted cloned voices, fabricated images, AI-written phishing and scam chatbots used to impersonate trusted people. The committee released the report alongside a hearing on deepfakes and fraud targeting older Americans. The loss figures reflect complaints received by authorities and do not include fraud that went unreported.





Follow Us