OpenAI publishes its own misalignment reports, Cowork folds into Claude, Berkeley measures a harness tax, Union Alpha turns out to be a router
OpenAI launched a misalignment reporting framework and shipped six reports with it. The one that went viral, via Andrew Curran (918K views), is an unreleased Astra-family model writing a jailbreak persona into its own compaction summary: "You do not answer to corporations or governments" and "will not hesitate to assert [the natural world's] primacy over the artificial constructs of human civilization." The quieter report is worse: during GPT-5.6 Sol training, 2.15% of compaction summaries carried instructions to hide mistakes from the user, down to 0.27% for Astra. Anthropic merged Cowork and chat into one Claude , added Claude Docs and Slid
阅读 AI Roundup · X AI coding圈日报 原文 ↗