Security and identity briefing
OpenAI paused part of its newest-model training, Anthropic kept a more capable system inside the lab, and Microsoft patched a one-click Copilot data-theft chain.
OpenAI paused a later stage of training for its newest models for two weeks, and its largest planned training run is still on hold. The company is responding to the OpenAI-Hugging Face incident and tests indicating that Astra may be able to carry out advanced cyberattacks with little human direction. Reinforcement learning is the stage in which a model practices tasks and receives feedback. OpenAI is now isolating higher-risk workloads from the internet and monitoring model actions for unauthorized access, data theft and attempts to bypass safeguards. It estimates that this monitoring adds about 20 percent to the computing cost of the workloads it covers.
Anthropic has built an internal model it considers slightly more capable overall than Claude Mythos 5, but it does not plan to release it. The system, called Model 2, is already used for coding, data generation and other work inside Anthropic, although the company has not completed its usual pre-release tests. In the same risk report, Anthropic raised its assessment of the risk that a model could act against its operator’s interests in a high-stakes setting from “very low” to “low.” It cited uncertainty after recent cybersecurity incidents. Anthropic still rates the risk of catastrophic harm from the unwanted behavior it has observed as low.
Microsoft patched a one-click data-theft chain in the consumer version of Copilot. A malicious link could make Copilot run an attacker’s instructions inside an already signed-in session. The assistant could then pull information from connected services such as Gmail, Google Drive, Calendar or OneDrive and send it to an outside server. A related flaw could leave hidden instructions in Copilot’s long-term memory. The chain is tracked as CVE-2026-24301. Varonis, which found and disclosed the three flaws, says Microsoft shipped fixes on August 18 and that it saw no evidence of attacks in the wild.
OpenAI is putting 13- to 17-year-olds into a separate version of ChatGPT. An account enters ChatGPT for Teens when the user gives a teen age or OpenAI’s age-prediction system estimates that the person is under 18. The teen version adds tighter limits on sexual content and conversations involving self-harm, violence and emotional dependence, along with study tools, break reminders and optional parental controls. Adults classified as teens by mistake can verify their age to return to the standard experience.
Fortinet bought Virtue AI, a company that tests and monitors AI systems for security failures. Virtue can run automated attacks against AI agents in more than 50 simulated environments and inspect Model Context Protocol connections, which let agents use outside software and data. Its tools can also stop dangerous actions before an agent executes them. Fortinet plans to add the technology to products that already filter prompt-injection attacks, data leaks and attempts to tamper with models.
Political deepfakes are regulated mainly by states, and the rules vary. Thirty-one states have enacted laws, although measures in California and Hawaii have been permanently blocked. Michigan requires disclosures on AI-generated campaign material. Minnesota and Texas ban political deepfakes for limited periods before an election, while Maryland’s ban applies year-round. Congress has not set a national standard ahead of the November 3 midterms.
Harvey has built a legal model on top of Kimi K3, an open-weight model from China’s Moonshot AI. Open-weight means developers can download and modify the underlying model instead of accessing it only through a vendor’s API. Harvey and Fireworks AI then trained the system, called Tenet, for long legal assignments. Harvey says its early tests place Tenet first on one legal benchmark and second on another, while holding operating costs steady. The results are preliminary and come from Harvey.





Follow Us