# Gemma 4 multi-turn greedy degeneration under a letter-level constraint

> Source: <https://discuss.huggingface.co/t/gemma-4-multi-turn-greedy-degeneration-under-a-letter-level-constraint/178281#post_2>
> Published: 2026-07-30 06:41:54+00:00

For now, I ran a small experiment in Colab (though **with Gemma 4 12B**):

I could not reproduce the same pathological collapse in the small reference control I tried.

The run used a Colab T4, Gemma 4 12B `Q4_K_M`

, and a current llama.cpp-derived CUDA path. I tested a hard “S” constraint with greedy decoding and a 300-token ceiling, a softer version of the constraint, the model-card sampling settings, two-turn histories, a short synthetic collapsed assistant history, and exact input lengths of 1023, 1024, 1025, and 1039 tokens. None of those cases entered the word-loop / mixed-script / symbol-debris state; the longest identical-token run was one in the recorded cases.

I would treat that only as a **negative proxy**, not as a reproduction or as evidence that the 31B path is clear. Google describes 12B as **12B Unified** and 31B as **31B Dense**; the model size, architecture, quantization, prompt, and runtime all differ, even though both use a 1024-token local sliding window ([Gemma 4 model card](https://ai.google.dev/gemma/docs/core/model_card_4)). The result mainly makes this narrower explanation less convincing:

greedy decoding + a prompt-level letter constraint is sufficient by itself to produce the published failure.

The strongest low-cost check I found is earlier in the stack. In the [incident snapshot of gemma4.rs](https://github.com/Ruffian-L/hydrodynamic-swarm/blob/97fec56/src/gemma4.rs), local-layer K/V is trimmed to the final 1024 positions

Replaying that arithmetic gives:

| Prefill length | Retained K/V | Query rows with no valid retained key |
|---|---|---|
| 1023 | 1023 | 0 |
| 1024 | 1024 | 0 |
| 1025 | 1024 | 1 |
| 1039 | 1024 | 15 |
| 2048 | 1024 | 1024 |

I may be reading the intended prefill contract incorrectly, but that geometry looks worth checking **before** interpreting the sampler ablations. It is a static code-level observation only: I have not established that the published run actually propagated NaNs or other non-finite values, or that this caused the transcript. The same trim-before-full-query-mask order also appears in the [current master implementation](https://github.com/Ruffian-L/hydrodynamic-swarm/blob/master/src/gemma4.rs), so the check is still relevant unless the executed build used another revision.

My default route would therefore be to add a small **A0 integrity gate** before A1–A6:

Two small reproducibility notes may also save time:

`chat_gemma4.sh`

`TOKENS="${100:-300}"`

. A command such as `./scripts/chat_gemma4.sh 48`

therefore still selects 300; only the 100th positional argument overrides it. The `${1:-80}`

. This does not explain the intentional 300-token incident, but it matters when interpreting any old “48-token” A1 result.So my current ordering would be:

The A1–A7 ladder in the post still looks useful. I would just put this small A0 gate in front of it, because it can distinguish a boundary or loader problem from a genuine generation basin before the later mitigations make both cases look “fixed.”
