# Harnesses, orchestration and multi-agent tools

> Source: <https://forum.level1techs.com/t/harnesses-orchestration-and-multi-agent-tools/257778#post_2>
> Published: 2026-10-07 13:58:19+00:00

I don’t think you’re misunderstanding the concepts. What you’re describing sounds more like a distributed multi-agent setup where the llama-server instances are the workers, and the missing piece is a communication/orchestration layer between them.

The JSON + `curl` approach makes sense as a starting point, but synchronous request/response will become a bottleneck once agents need to coordinate with several peers or work asynchronously. I’d probably separate the model-serving layer from the messaging layer rather than expecting llama.cpp itself to handle all of the orchestration.

For example, the llama-server instances could remain responsible for inference, while a lightweight message broker or event-based layer handles agent-to-agent communication, task routing, retries, and authentication. That would also give you more flexibility to change the orchestration framework later without rebuilding the model-serving side.

The TLS requirement is definitely reasonable too. I’d rather solve the certificate-handling problem in the communication layer than fall back to plaintext just because the orchestration tool has limited support.

The interesting part of your setup is that the agents are already distributed across dedicated LAN machines. At that point, you’re essentially dealing with a small distributed system, so reliability, authentication, timeouts, and asynchronous messaging may matter more than which framework happens to call itself an “agent harness.”
