Latent-GRPO: Reinforcement Learning in Continuous Thought Space Latent-GRPO, a reinforcement-learning method that replaces discrete text tokens with continuous recurrent thought vectors in embedding space, eliminates the token tax that causes 19% to 34% of generated trajectories in gft-studio multi-hop routing runs to hit the sequence ceiling mid-sentence and receive zero reward. The method feeds internal recurrent representations directly back into the model in continuous space (d = 5,120) instead of projecting hidden states onto a 152,064-token vocabulary, addressing the bottleneck that forces models such as Qwen3.6-27B through 64 transformer layers per intermediate thought word. The approach targets Group Relative Policy Optimization (GRPO), where truncated completions fail their verifier and poison gradient updates. When you sit down to solve a complex puzzle or plan three moves ahead in chess, do you narrate every synaptic firing to yourself in full, grammatically correct English sentences? Of course not. Human cognition operates across continuous, multi-dimensional mental representations—spatial, relational, intuitive. We only convert our thoughts into human language when we need to speak to someone else. Yet the dominant paradigm in AI reasoning today from DeepSeek-R1 to OpenAI o1 forces neural networks to do the exact opposite. We train models to emit thousands of discrete text tokens enclosed in