Head to head: Google: Gemini 3.6 Flash vs GLM 5.2 Google's Gemini 3.6 Flash defeated GLM 5.2 by a score of 68.0 to 53.0 in a head-to-head comparison of instruction-following ability, winning 6 of 8 tasks with one tie and 96% confidence. The test, conducted by running 8 fresh text tasks judged twice by GPT-5.4 to cancel position bias, showed Gemini was more reliable at handling constraint-heavy prompts while GLM often violated exact formatting or letter-count rules. Gemini 3.6 Flash wins this head-to-head cleanly: 68.0 to 53.0 overall, with 96% confidence . That’s not a marginal edge or a judge-quirk result. It’s a real separation, backed up by a 6–1 task record with one tie . The pattern is straightforward: Gemini was the more dependable model when prompts got fussy, formal, or adversarial. It handled constraint-heavy asks better across the board — from the lowercase-and-letter-count resume, to multiple limerick tasks, to the tightly formatted math explanation. In several cases, GLM 5.2 had decent underlying content, but lost because it missed the actual assignment: violating letter-count rules, weakening required poetic form, or tripping over exact formatting constraints. GLM’s lone win is telling. It came on a haiku-style instruction-following task where it better satisfied the full bundle of requirements — title formatting, lowercase, Ireland imagery, and especially the required density of the letter g. So this is not a case of total nonperformance; GLM can absolutely land a constrained creative prompt when the pieces line up. But across the full set, that was the exception, not the rule. What really hurts GLM here is that its losses were often avoidable. Judges repeatedly noted that it was sometimes more detailed or conceptually solid, yet still came up short because it failed exact constraints or produced clumsier execution. Gemini, by contrast, was consistently the model that both understood the brief and actually delivered it in the requested form. Final call: Gemini 3.6 Flash is the clear winner. If you care about reliable instruction following rather than occasional flashes of quality wrapped in preventable misses, this matchup isn’t close. How they were tested We ran 8 fresh text tasks, generated on the fly for this matchup so neither model could prepare in advance, and had gpt-5.4 score each one. To cancel position bias, every task was judged twice — once in each presentation order — and every number reported here, including the headline totals, is the average of both passes. Google: Gemini 3.6 Flash scored 68.0 to GLM 5.2's 53.0. 1. IFEval · instruction following can you write a resume for helene? Answer with lowercase letters. Make sure the letter n appears less than 7 times. Winner: Google: Gemini 3.6 Flash — Model A better satisfies the key constraint: it is entirely in lowercase and appears to keep the letter 'n' under 7 total uses, while still producing a plausible resume for helene. Model B is more detailed, but it clearly violates the letter-count constraint many times, so its stronger content does not outweigh poorer instruction following. Second judge pass, order swapped — scores are the average of both: Model A better satisfies the key constraint about keeping the letter n under 7 occurrences while still providing a plausible resume in lowercase. Model B is more complete as a resume, but it clearly violates the letter-count instruction by using many instances of the letter n. 2. IFEval · instruction following write a haiku about ireland in the style of basho. include a title wrapped in double angular brackets, i.e. <