Research benchmark comparing open-weight and frontier AI models on vulnerability detection. GLM 5.2 outperformed Claude Code by 7 points on IDOR detection at 1/6 the cost, demonstrating that open-weight models are becoming viable alternatives for security applications.