Anthropic red team reports GLM-5.3 control flow hijacks in binary exploitation testsMachine translation
Anthropic Frontier Red Team evaluated several models on 100 randomly selected tasks from its internal Binary Exploitation benchmark. GLM-5.3 developed full control flow hijacks in 4% of trials, versus 6% for Claude Mythos Preview; earlier models Claude Opus 4.6 and GLM-5.2 succeeded in none of those tasks.
Anthropic red team reports GLM-5.3 control flow hijacks in binary exploitation tests
Anthropic Frontier Red Team evaluated several models on 100 randomly selected tasks from its internal Binary Exploitation benchmark. GLM-5.3 developed full control flow hijacks in 4% of trials, versus 6% for Claude Mythos Preview; earlier models Claude Opus 4.6 and GLM-5.2 succeeded in none of those tasks.
材料 38c902b0bf1a4fcb82cece06b6967487;建议 5a6e7b960d034613bce238b14d86d8a8;系统证据核验通过,非人工审稿。
Read at the original source