cd /news/ai-infrastructure/the-private-ai-stack-is-the-product · home › topics › ai-infrastructure › article
[ARTICLE · art-141662] src=lm-kit.com ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

The Private AI Stack Is the Product

LM-Kit argues that private AI deployments should be bought as one integrated engineering system rather than assembled from separate components, citing its LM-Kit.NET SDK and LM-Kit One private AI server as sharing one runtime, one scheduler, one memory model and one release. The company reports that replacing OpenSearch clusters behind its group's products with LM-Kit One's built-in search divided the cost of search by 10 while improving both accuracy and latency. LM-Kit also describes a throughput optimization that moved prefill of image-bearing prompts to a side context, raising throughput but lowering accuracy on fine-tuned models because the side context ran on base weights instead of the customer's adapter, a regression now pinned by an end-to-end test.

by read5 min views1 publishedSep 28, 2026
The Private AI Stack Is the Product
Image: Lm-Kit (auto-discovered)

A team decides to bring AI in-house. Six months later they run a Python inference server, a separate embedding service, a vector database, a document parser, an OCR engine, an orchestration framework, a Node.js gateway and a monitoring stack. Each one was a sound choice. Each one ships on its own schedule, carries its own security advisories and expects its own specialist. The team set out to use AI. It now runs an AI infrastructure company on the side.

We see versions of this situation all the time. These teams can assemble the stack; that was never the hard part. The hard part is keeping it good while every layer underneath it changes.

You get the system, not the best component #

A production AI stack does not perform at the level of its best part. It performs at the level of everything working together, and that is where assembled stacks break down.

        Local optimization does not guarantee global optimization.

A new embedding model raises retrieval scores and degrades one production workload. A new parser extracts tables better and segments pages worse. A faster inference path changes numerical behavior. Dropping a preprocessing step cuts latency and costs accuracy three stages later. Each change looks like progress where it happens. Only the end-to-end outcome says whether it was.

When the layers come from different projects, no individual project owns the end-to-end outcome. Each project optimizes what it can see.

What integration looks like in practice #

LM-Kit is one engineering system: the LM-Kit.NET SDK and LM-Kit One, the private AI server built on it. Inference, embeddings, document parsing, OCR, retrieval, search, speech, agents and the security controls around them share one runtime, one scheduler, one memory model and one release. Three examples show why that matters.

  • One improvement, every workload. Text generation, vision encoding, OCR, layout detection and speech recognition all draw CPU threads from one process-wide budget. When we improved how that budget is shared, every one of those workloads gained at once. An assembled stack cannot do this: each runtime assumes it owns the machine, and together they oversubscribe it.
  • Faster, and quietly wrong. To raise throughput, we moved the prefill of image-bearing prompts onto a side context that runs while other requests keep generating. Throughput went up. Accuracy on fine-tuned models went down, because the side context ran on the base weights instead of the customer's adapter. No component benchmark would have caught it. The end-to-end extraction tests did, and a regression test now pins the fix. This is the gap between local and global optimization, in one change.
  • The same optimization, two answers. One speculative decoding strategy speeds up generation on one class of GPU and slows it down on another. So the engine does not pick a winner once; it prices each option against measured cost on the hardware it is actually running on, and chooses per pass.

Integration compounds the same way at the product level. When we replaced the OpenSearch clusters behind our group's products with LM-Kit One's built-in search, the migration divided the cost of search by 10 while improving both accuracy and latency, and one whole external system left the architecture.

Engineering by measurement #

None of this works on intuition. Our loop is plain: observe, isolate, experiment, benchmark, validate, integrate, then measure again at the level of the complete system.

Two rules decide whether a change ships. It must generalize beyond the cases that inspired it. And it must improve something that matters, whether speed, accuracy, memory or robustness, without an unacceptable regression elsewhere. Native engine changes run the full test suite on GPU hardware before they are committed. Fine-tuning changes run a training regression battery against banked results. A change that moves a banked number on purpose has to say why.

The measurement system is never finished either. Bad measurements produce bad engineering decisions, so the benchmarks get the same scrutiny as the code they judge.

We use AI heavily in this loop. It writes experiments, bisects regressions and explores alternatives a small team could never afford to try by hand. It widens the search. The benchmark still decides, and a human still decides what earns a place in a system meant to last for years.

Integrated does not mean closed #

The fair objection to all this: haven't you just swapped ten open-source black boxes for one vendor black box? No, and we design against it.

  • Your hardware, your perimeter. Everything runs on your machines, on Windows, Linux or macOS, on CPU or GPU. No data crosses a boundary you did not choose.
  • Your models. Use the catalog, bring your own weights, or fine-tune on your own data.
  • Standard interfaces. LM-Kit One speaks the OpenAI- and Ollama-compatible APIs and the Model Context Protocol, so existing clients and tools connect without rework.
  • Open source inside, understood deeply. We build on projects such as llama.cpp, and keep our changes to them as a tracked patch series that follows upstream. When a component matters to the stack, we learn it well enough to measure it, modify it and replace it when something better appears.

The aim was never for every internal component to beat every alternative forever. The aim is for the whole system to keep improving, and to adopt outside innovation whenever the numbers show that it improves the system.

What compounds over time #

Building an impressive private AI stack once is within reach of any capable team. Keeping it sharp for five years, while models, runtimes, hardware and threats keep changing, is a different job, and it grows faster than most teams expect.

What compounds is not a feature. Features can be copied. What compounds is the machinery that turns new research, open-source progress and real customer workloads into measured improvements across the whole stack, release after release: the shared runtime, the benchmarks, the accumulated optimizations and the understanding of how the layers interact.

Companies should not have to become AI infrastructure companies because they need private AI. Carrying that complexity is our job.

The private AI stack is the product. We build it as one system, and we measure it as one.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @lm-kit 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-private-ai-stack…] indexed:0 read:5min 2026-09-28 · —