Simon WillisonVoices
Quoting Anthropic Frontier Red Team
We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude…
Read at Simon Willison ↗Related

VoicesBREAKING: Secret US AI evaluation framework has been partly revealed Gary Marcus

ValleyOpenAI hit with landmark lawsuit following Hugging Face hack Axios Technology

ResearchHow Diffusion Controller unifies and simplifies AI image generation Google Research

SiliconQualcomm’s System Level Architecture in the Snapdragon X2 Elite Chips and Cheese

GitHubDeveloper policy update: Transparency, state policy, and what’s ahead GitHub Blog
