{"slug": "turnstone-orchestrating-ai-coding-agents-across-local-hardware", "title": "Turnstone: Orchestrating AI Coding Agents Across Local Hardware", "summary": "Turnstone is an orchestration layer that distributes AI coding agent tasks across multiple local machines, including DGX Station, GB10 (DGX Spark class) clusters, and RTX Pro 6000 rigs, rather than routing every prompt through a single chat session. The system ships with built-in personas — engineer, executive, manager, orchestrator, researcher, scribe, and writer — and uses a manager/orchestrator persona that spawns workers plus an evaluator to split a project into subtasks, execute them, and verify results before marking them done. Turnstone's stated advantage is tiered model usage, letting a strong model such as Claude Opus act as judge or orchestrator while cheaper local models (demonstrated with Nemotron and DeepSeek variants) handle token-heavy work like unit and integration tests; nodes can be Docker containers or full machines, and a worker node needs no GPU of its own if pointed at a model served elsewhere.", "body_md": "# Turnstone: Orchestrating AI Coding Agents Across Local Hardware\n\nTurnstone lets you distribute AI coding agent tasks across DGX Station, GB10 clusters, and RTX Pro 6000 rigs. Here's how it actually works.\n\n## What is Turnstone?\n\nTurnstone is an orchestration layer for running AI coding agent tasks across multiple local machines instead of a single chat window. Rather than sending every prompt to one model, it lets you register several machines as nodes, assign different models and personas to them, and hand off whole development tasks (fix a bug, add a feature, port a repo to a new architecture) to a manager process that breaks the work into subtasks and schedules them on whatever hardware is available.\n\n## TL;DR\n\n- **Turnstone distributes coding tasks across nodes** instead of running everything through one chat session, letting you point it at a repo and have it schedule work across whatever local machines you’ve registered.\n- **It ships with built-in personas** like engineer, executive, manager, orchestrator, researcher, scribe, and writer, each tunable with its own instructions and memory, and some support MCP (model context protocol) servers for extra tooling.\n- **A manager/orchestrator persona spawns workers and an evaluator** for a given task, so a project gets split into subtasks, worked on, and then checked before being marked done.\n- **Hardware choice changes performance meaningfully** , with a DGX Station’s HBM3e memory crushing prompt processing and single-user speed, while a cluster of RTX Pro 6000s or GB10 (DGX Spark class) machines can hold larger models and serve more concurrent users.\n- **The real advantage is tiered model usage** : a strong model like Claude Opus can act as a judge or orchestrator while cheaper, local models (demonstrated with Nemotron and DeepSeek variants) do the bulk of the token-heavy grunt work like unit tests and integration tests.\n- **Nodes can be Docker containers or full machines** , including DGX Spark units run directly off their shell for full hardware access, and a worker node doesn’t need its own GPU if it’s pointed at a model served elsewhere.\n\n## Seven tools to build an app. Or just Remy.\n\nEditor, preview, AI agents, deploy — all in one tab. Nothing to install.\n\n## How does Turnstone distribute work across machines?\n\nYou start by registering a model endpoint, typically a base URL and IP address for something running locally. Once a model is connected and enabled, you can add more, mixing models that are good at different things. Turnstone also lets you configure routing rules in general terms (certain kinds of requests go to certain models), and you can build custom personas beyond the defaults it ships with.\n\nWhen you give it a task, like cloning a repository and setting up instructions on a specific machine, Turnstone’s orchestrator persona decides where to run it. In the workflow shown, a user had a repository of bootstrap scripts for setting up developer machines (Windows, Linux, Mac) and asked Turnstone to create an ARM-compatible branch, since the existing scripts assumed x86 and broke on ARM-based systems for tools like Python and Node. The orchestrator spun up workers, including an auditor to check the output, and scheduled the work on an available node automatically. A second task, finding and fixing a specific bug in a separate repository, ran at the same time on the same machine, showing that Turnstone can parallelize multiple independent jobs rather than handling one request at a time.\n\n## What are nodes, personas, and the manager/evaluator pattern?\n\nNodes in Turnstone are the individual compute environments it can dispatch work to. They can be Docker containers with limited access until you mount a shared workspace (by default at `/workspace`), or they can be full machines accessed more directly, such as running Turnstone straight from the shell on a DGX Spark, which gives it real hardware access rather than a sandboxed container.\n\nPersonas define the role each worker plays: engineer, executive, manager, orchestrator, researcher, scribe, and writer were the defaults mentioned. Each can be customized, down to specific writing preferences (the scribe persona can be told to avoid certain punctuation habits, for instance), and Turnstone supports saved memories so instructions persist across sessions instead of being repeated every time.\n\nFor a coding task, the orchestrator creates a manager and an evaluator alongside worker agents. The manager breaks the job into subtasks and tracks them; the evaluator checks the work before it’s considered complete. In one demonstrated run, a bug-fix task completed using around 15,000 tokens, with a breakdown available showing what each subtask consumed, giving fairly granular auditability into what the system actually did.\n\n## Does the hardware you run it on matter?\n\nYes, significantly, and the differences come down to memory architecture rather than just raw GPU count.\n\nA DGX Station has 252 GB of HBM3e memory, which moves data at roughly 7 terabytes per second, an order of magnitude above GDDR7 on consumer and workstation cards. The station also has a total of 748 GB of unified memory, so larger models that spill past the HBM pool still run, just slower, off LPDDR5 at around 600 GB/s. That architecture pairs well with newer mixture-of-experts models, where some weights are fine running from slower memory while active parameters stay in HBM.\n\n## Remy doesn't write the code. It manages the agents who do.\n\nRemy runs the project. The specialists do the work. You work with the PM, not the implementers.\n\nA setup using eight RTX Pro 6000 GPUs takes a different path: those cards communicate over PCIe Gen 5 at roughly 64 GB/s per direction per link, far slower than HBM, but the aggregate VRAM across eight cards can be large enough to hold models that wouldn’t fit in a DGX Station’s HBM pool outright. In testing with a DeepSeek V4.1 variant at FP8 (around 614 GB in size), the RTX Pro 6000 setup could fit the full model, while a GB10-based “Spark” style cluster running a similar model via DSpark hit around 307 tokens per second, edging out the HBM-constrained station by a small margin on that particular workload, despite the station’s faster memory bandwidth in other scenarios.\n\nPrompt processing (prefill) consistently favored the HBM-equipped DGX Station by a wide margin, since nothing beats HBM for raw prefill speed. For single-user generation, a 120B-parameter open model (GPT-OSS 120B) ran entirely from HBM on the station at notably high speed.\n\n## Why run models locally instead of just using cloud APIs?\n\nMulti-user throughput is the practical argument. If you’re buying hardware like this, it’s usually to support a development team or to keep work private rather than to beat cloud pricing per token, since cloud tokens remain comparatively cheap against the cost of owning one of these machines. The economics only make sense if the hardware stays busy continuously.\n\nIn throughput testing across multiple simultaneous users, the hardware handled around 3,800 tokens per second in aggregate, and a Nemotron 3 Super (12B) model hit just over 2,000 tokens per second under similar multi-user load. Models in that size range aren’t frontier-level, but they’re capable enough for real development work: unit tests, integration tests, and routine coding tasks with reasonable context.\n\nThe pattern that emerges is tiering: a strong model like Claude Opus can run as a judge or orchestrator, spending tokens at a slower rate while directing eight or ten local worker agents that do the token-heavy labor without touching cloud API costs. The local models don’t need to be as capable as the frontier ones, just good enough for well-scoped tasks under supervision.\n\n## Is Turnstone worth setting up?\n\nFor someone who already owns or has access to more than one AI-capable machine, such as a DGX Station, a DGX Spark cluster, or several RTX Pro 6000 workstations, Turnstone’s value is in turning separate boxes into one schedulable pool instead of manually SSHing into each one to kick off a coding session. The personas, memory, and evaluator pattern add structure that a single raw chat session with a coding model doesn’t give you by default, and the auditability (seeing subtasks and their token usage) matters for anyone running agents somewhat unsupervised.\n\nFor a single-machine, single-user setup, the orchestration overhead is probably unnecessary. The approach clearly targets small teams or privacy-sensitive workloads where keeping expensive hardware continuously busy, and keeping code off third-party servers, outweighs the convenience of just paying for cloud tokens.\n\n## Frequently Asked Questions\n\n### What hardware did the demonstration use?\n\nA DGX Station (GB300-based, with 252 GB of HBM3e and 748 GB total unified memory), a machine with eight RTX Pro 6000 GPUs connected via PCIe Gen 5, and a cluster of DGX Spark (GB10) units.\n\n### Can Turnstone run on a mix of different machine types?\n\nYes. Nodes can be Macs, Linux boxes, DGX Spark units, or Docker containers, and a worker node doesn’t need its own GPU if the model it’s using is served from elsewhere on the network.\n\n### What models were used in testing?\n\nExamples mentioned include GPT-OSS 120B, DeepSeek V4.1 (including a “Flash” variant), Nemotron 3 Super (12B), and smaller DeepSeek 0731 variants used as workers or judge models.\n\n### What is the manager/evaluator pattern?\n\n## \nPlans first.\n*Then code.*\n\nRemy writes the spec, manages the build, and ships the app.\n\nWhen Turnstone’s orchestrator takes on a task, it spins up a manager to break the work into subtasks and an evaluator (auditor) to check the output before the task is considered complete.\n\n### Does a bigger model always perform better on this hardware?\n\nNot necessarily. Model architecture matters as much as size. Mixture-of-experts models can run efficiently even on slower unified memory because not all their weights need to sit in high-bandwidth memory at once, which changes how you should match a model to a given machine.", "url": "https://wpnews.pro/news/turnstone-orchestrating-ai-coding-agents-across-local-hardware", "canonical_source": "https://www.mindstudio.ai/blog/turnstone-ai-orchestration-local-hardware/", "published_at": "2026-10-08 00:00:00+00:00", "updated_at": "2026-10-08 10:17:43.254723+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "agent-protocols", "ai-infrastructure", "developer-tools"], "entities": ["Turnstone", "DGX Station", "GB10", "RTX Pro 6000", "DGX Spark", "Claude Opus", "Nemotron", "DeepSeek"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/turnstone-orchestrating-ai-coding-agents-across-local-hardware", "markdown": "https://wpnews.pro/news/turnstone-orchestrating-ai-coding-agents-across-local-hardware.md", "text": "https://wpnews.pro/news/turnstone-orchestrating-ai-coding-agents-across-local-hardware.txt", "jsonld": "https://wpnews.pro/news/turnstone-orchestrating-ai-coding-agents-across-local-hardware.jsonld"}}