Z.ai said ⁠GLM-5.3 scored 84.5% on CyberGym, a test of whether a model can review code, identify security flaws and confirm that they are real