Parsing JSON from Thinking-Model APIs A developer has documented a defensive parsing strategy for handling JSON responses from reasoning-model APIs, which often split structured output across multiple content parts, prepend conversational preambles, and attach internal reasoning traces. The approach concatenates all text parts, strips markdown code fences, attempts a strict JSON.parse, and falls back to a regex that isolates the outermost curly braces. The writeup also warns that reasoning models bill internal monologue against the same token limit as the final answer, so output budgets must be scaled up to avoid truncation. When you request structured data from a reasoning model using a JSON mime type, the response often arrives split across several content parts rather than sitting neatly in a single text block. This breaks standard assumptions in ingestion pipelines. Developers usually expect the model to return a single text payload starting with an opening brace. Reasoning models, however, routinely prepend prose preambles, attach internal thought signatures, and partition the final output across distinct blocks. For a related implementation, see Schema First Gates Ai Publishing Pipelines https://raylabs.app/articles/schema-first-gates-for-ai-publishing-pipelines/ . The core problem manifests in three distinct ways during API integration. First, if your code attempts to read only the first part of the response, it captures conversational filler like introductory remarks, causing JSON parsers to throw syntax errors. Second, the model's token output limit must accommodate both the internal reasoning trace and the final JSON payload. If you allocate a small budget, the model spends all its tokens on thinking, leading to abrupt truncation with a max tokens finish reason. Third, applying blind retries to these truncated payloads merely drains rate limits without resolving the underlying deterministic failure. A typical API payload from a reasoning model arrives as an array of content parts. The first part may contain a conversational preamble. Subsequent parts can include encrypted or signed reasoning traces, followed eventually by the markdown-fenced JSON structure. Treating this stream as a monolithic document guarantees parser failure. Consider a scenario where an agent requests a configuration object from a reasoning model with an output budget set to two hundred tokens. The model generates an extensive internal monologue to work through the schema requirements. It exhausts the token limit just as it begins writing the payload, returning a truncated string without a closing brace. Your application receives a finish reason indicating truncation, while the primary text parts contain only partial data and reasoning debris. To handle multi-part responses reliably, your ingestion pipeline needs a defensive parsing strategy that combines concatenation, unfencing, strict parsing, and robust fallbacks. For a related implementation, see Multi Agent Review Pipeline https://raylabs.app/articles/multi-agent-review-pipeline/ . interface ModelPart { text?: string; } interface ModelResponse { candidates?: Array<{ content?: { parts?: ModelPart ; }; finishReason?: string; } ; } function extractJsonPayload response: ModelResponse : Record