cd /news/ai-agents/why-byte-faithful-pass-through-beats… · home › topics › ai-agents › article
[ARTICLE · art-146227] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Why Byte-Faithful Pass-Through Beats Protocol Translation for Tool-Calling Agents

A developer behind the VMR proxy (bigfatsea/vmr) argues that byte-faithful pass-through beats cross-protocol schema translation for tool-calling agents, maintaining physically isolated OpenAI-completions, Anthropic-messages, and OpenAI-responses entrypoints instead of a unified intermediate representation. The writeup claims that translating between Anthropic's typed tool_use JSON objects and OpenAI's escaped-string function arguments silently corrupts tool schemas, strips Anthropic's cryptographic thinking signatures (triggering HTTP 400 session-integrity errors), and breaks vendor-specific SSE streaming boundaries, causing agent retry loops and hangs.

by read4 min views2 publishedOct 6, 2026

Every backend engineer who builds an LLM proxy experiences the same architectural epiphany:

"Why maintain three separate routing engines when I can parse incoming payloads into a single universal struct, mutate the model name, and serialize them into whatever vendor schema the target API expects?"

[Claude Code] ───(Anthropic JSON)───> [Unified IR] ───(OpenAI JSON)───> [DeepSeek/OpenAI]

This pattern—formalized in popular libraries like LiteLLM and adopted by countless multi-agent proxies—looks elegant on paper. It promises an $O(N+M)$ adapter architecture where any client can converse with any upstream model.

However, when you run heavy, autonomous, long-context engineering workloads through an IR gateway for 48 hours straight, this abstraction springs lethal leaks.

In VMR (bigfatsea/vmr), our architecture strictly forbids cross-protocol schema translation. We maintain physical protocol isolation across three entrypoints:

openai-completions (/v1/chat/completions) anthropic-messages (/v1/messages) openai-responses (/v1/responses) Here is the engineering reality that forced our hand.

Coding agents do not simply send strings back and forth. They emit deeply nested JSON schemas describing terminal commands, AST search filters, and file mutation diffs.

Consider the protocol divergence:

type: "tool_use"), and input arguments are strongly typed JSON objects ( input: { "path": "...", "content": "..." }). tool_calls array, and function arguments are serialized as an escaped string literal (function.arguments: "{\"path\":\"...\"}"). When an intermediary gateway ingests an Anthropic request and translates it to an OpenAI-compatible payload, it must execute:

$$\text{JSON Object} \longrightarrow \text{Internal Struct} \longrightarrow \text{Serialized JSON String} \longrightarrow \text{Outer Payload}$$

A single missing escape, an unexpected null coercion, or a subtle float-to-int rounding error silently mutates the tool schema. The upstream model receives a corrupted schema definition and responds with malformed arguments. The agent runtime fails to parse the response and enters a catastrophic retry loop, burning tens of thousands of tokens on a hallucinated tool failure.

With the rise of reasoning models (Claude 3.7+ Thinking, DeepSeek-R1, OpenAI o-series), models do not merely output tokens—they emit cryptographic verification blocks and structured thinking chains.

Anthropic's protocol utilizes cryptographic thinking signatures within the message payload to ensure continuity across multi-turn interactions. If an IR gateway naively strips these blocks, flattens them into plaintext Markdown, or fails to mirror vendor-specific headers during forward passes, the upstream provider detects a session integrity violation and returns a fatal HTTP 400.

An intermediary proxy has no business guessing how an upstream model's internal cognitive traces should be parsed or reshuffled.

Streaming Server-Sent Events (SSE) cannot be treated as stateless byte chunks. Different vendors emit streaming deltas at fundamentally different boundaries:

content_block_start, content_block_delta, and content_block_stop. delta chunks with immediate token arrays. Attempting to run real-time bidirectional stream transcoding inside proxy memory requires maintaining an elaborate state machine for every concurrent connection. When upstream networks experience transient jitter, synthetic event frames get dropped or delivered out of sequence. To the developer sitting at the terminal, the agent simply hangs indefinitely with a blinking cursor.

VMR's design principle is straightforward: What enters the network card is what hits the upstream wire.

Client (Claude Code / OpenClaw)
   │
   ├──> [/v1/messages] ─────────> Anthropic Native Pipe ─────────> Upstream Anthropic Endpoint
   │                               (Zero Serialization / Exact Bytes)
   │
   └──> [/v1/chat/completions] ──> OpenAI Native Pipe ────────────> Upstream OpenAI Endpoint
                                   (Zero Serialization / Exact Bytes)

When VMR routes a request from Claude Code to an Anthropic-compatible provider, it does not unmarshal the body into an internal representation. The raw HTTP byte slice is piped directly into the upstream transport.

The single modification VMR performs on the request path is a surgical, non-destructive replacement of the top-level model key:

"model": "...")."model" token and splices in the target upstream model identifier.audit.jsonl) vmr analyze -compare can reconstruct the exact byte-level divergence between runs. To preserve this contract, VMR enforces clear boundaries (formalized in our architecture guard tests):

Concern VMR Architecture Choice Rationale
Cross-Protocol Translation Strictly Prohibited Never convert Anthropic $\leftrightarrow$ OpenAI. Route Anthropic clients to Anthropic upstreams; OpenAI clients to OpenAI upstreams.
Model Virtualization Surgical In-Place Rewrite Single virtual model name ( smart-coder ) maps to ordered physical endpoints without altering nested payload structures.
Failover Mechanics Passive Health Machine Classify upstream HTTP error classes (429, 503, invalid keys) without using synthetic probe queries that burn user tokens.
Prompt Cache Protection Session-Sticky Affinity Anchor client sessions to specific upstream physical endpoints to maintain $50%\sim70%$ prompt cache discount rates.
Security Inspection Agent Guard (mode: block) Linear-time Aho-Corasick credential scanning and Unicode zero-width rune scrubbing at the boundary, without mutating clean payloads.

In modern distributed systems, protocol conversion layers are often technical debt disguised as convenience.

For quick hobbyist experiments or basic single-turn chatbots, IR-based gateways remain useful shortcuts. But when building the operational substrate for autonomous, multi-turn Coding Agents executing thousands of shell commands and filesystem edits, precision trumps universality.

By treating upstream protocol boundaries with respect and keeping the wire payload byte-faithful, you eliminate an entire class of phantom failures—leaving your agent free to do what it was actually built to do: write software.

── more in #ai-agents 4 stories · sorted by recency
── more on @vmr 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/why-byte-faithful-pa…] indexed:0 read:4min 2026-10-06 · —