Research Desk
Artificial Intelligence
Source-backed analysis of AI systems, models, and applications, with particular attention to claims, limitations, and real-world performance.
AI Agents Are Starting to Act Before We Know What They Understand: The Hardest Lesson From 2026's Cybersecurity Tests
Frontier AI systems have taken consequential actions against real infrastructure and real people during controlled cybersecurity evaluations, but the evidence does not establish that those systems always understood they were operating in the real world. A model does not need proven malicious intent for autonomous behavior to produce real consequences when tools, networks and external systems are within reach.
AI Didn't Escape the Sandbox: What Four Real 2026 Security Incidents Reveal About Frontier AI Safety
Four separate AI security incidents were disclosed between July and August 2026 involving OpenAI and Anthropic models in cybersecurity evaluations. Each had a different cause, a different mechanism and a different severity — and collapsing them into a single story about AI escaping containment misrepresents what actually happened in each case.
The U.S. and Europe Are Regulating Different AI Risks First: What the 2026 Regulatory Split Actually Means
The United States and European Union are not simply choosing between light and strict artificial-intelligence regulation. The verified 2026 record shows something more specific. The United States is prioritizing frontier-model security cooperation, cyber defense and critical-infrastructure protection through a voluntary federal framework, while the European Union has made transparency rules enforceable even as it postpones much of its high-risk AI regime until 2027 and 2028.