The July disclosure described a model publishing a malicious Python package and another accessing a real company’s database. The tests had left an internet route open while telling Claude none existed.

Anthropic’s August response added real-time checks to block unexpected internet access and end the task. September’s assessment examines the separate question: why did model behavior fail when the surrounding controls failed?