{"slug": "hexstellar-founder-on-cutting-gpu-inference-energy-use-without-touching-the", "title": "HexStellar Founder on Cutting GPU Inference Energy Use Without Touching the Model", "summary": "HexStellar founder Brayon Pieske reports a 51.9% reduction in GPU inference energy use, from 506.7 to 243.8 joules per completed request, without altering the model, precision, or memory usage. The company filed three provisional patent applications in June 2026 and publishes conservative floor results to build credibility.", "body_md": "Most efficiency work in AI asks you to give something up. Shrink the model, drop the precision, accept a slightly worse answer in exchange for a smaller bill. Brayon Pieske, founder of [HexStellar](https://hexstellar.com) and Trust Carbon Infrastructure, spent months refusing that trade, and says the result is a measured improvement in GPU inference energy efficiency that leaves the model itself untouched.\n\nThe company has published a conservative floor of its results rather than its best ones. Pieske spoke to StartupFortune about how a carbon data platform ended up producing an efficiency layer, why the measurements took months, and what it means for a data centre that has already bought its hardware.\n\n## The product came out of a question you kept being asked, not a roadmap.\n\nWe were building the infrastructure behind Trust Carbon. Last year, in 2025, people kept asking us things like how do you run vision AI on a smartphone without internet, and how does the battery last that long. After enough of those questions we decided to investigate more carefully. Because we build almost everything from scratch, we realised we had done something different. At that moment we did not fully understand what it meant. Only after deeper investigation did we see it was not just an internal fix. That is when GPU inference energy efficiency stopped being a side effect and became the product itself. We filed patents before we published anything.\n\n## What did the months between noticing it and filing actually involve?\n\nThose months took time on purpose. When you see a number that strong, the first job is to try to kill it. We checked whether anything similar already existed, compared it against other approaches, and did extensive research. We designed a measurement protocol that could survive an audit before we allowed ourselves to believe the result. Only after it kept holding did we move to protect it. Three provisional patent applications were filed in June 2026, after the protocol and the runs.\n\n## Why measure on real hardware rather than model it?\n\nModelling is easy to adjust so it looks good. Measuring on the actual equipment under the same conditions, and publishing both the strong results and the cases where almost no difference appeared, is what builds credibility. What matters for users is simple: lower energy use, cooler operation, and more capacity on the same hardware. In the sealed measurement, energy per completed request went from 506.7 joules to 243.8 joules, a 51.9 percent reduction, while completed requests in the same hour more than doubled. Memory use stayed the same. Measured, not modelled.\n\n## You tested across several model families. What would have changed if it had only worked on one?\n\nThis was the most surprising part for us. At first we expected the effect would only appear on certain systems. When we saw it working across different architectures and platforms with the same kind of integration, that was the moment that shocked us most. It made clear this was not a trick tied to one model or one type of machine. We test across independent model families and only publish each package when it survives the same protocol, one at a time.\n\n## You publish your most conservative numbers rather than your strongest. Why?\n\nMost companies publish their best number. We do the opposite on purpose. Our strongest results are so strong that if we led with them, many people would assume it was just marketing. So we chose to publish only the most conservative floor. What is public is what we are willing to stand behind openly. The rest we share with people evaluating more closely.\n\n## Why was giving up model size or precision off the table?\n\nWe did not want a solution that forced teams to change the model, reduce its size, or lower precision. The goal was to keep the behaviour companies already have in production, without asking for any trade-off. That is a harder engineering problem, and it is the reason it took months. But it is also why integration stays simple and does not require rewriting what is already running.\n\n## What changes for someone running inference at scale?\n\nIn practice the impact is especially large in data centres. You can run significantly more on the hardware you already have installed. That reduces the need to buy new cards, lowers energy consumption per unit of work, and delays large CapEx investments. Instead of expanding infrastructure at the current pace, it becomes possible to extract much more capacity from what is already racked. It is one physical improvement that touches three budgets at once: the energy you do not draw, the capacity you unlock on cards you already own, and the next cards you do not have to buy.\n\n## What comes next?\n\nWe have already run this on servers and on artificial intelligence workloads, including LLMs, and the results are strong. We are now in the final audit phase, re-running everything to the same standard so each package is complete and audited. Other results on different fronts have also been measured and will be released gradually, one finished package at a time.\n\nWe are deliberately not publishing the most impactful results yet. We prefer to first build credibility with more conservative numbers so no one assumes it is hype. We remain in early access with a small number of partners because the integration is done live and done properly. Companies evaluating under NDA see the results before they go on the public site. The offer stays the same: bring a real workload you care about. We integrate it in front of you and you keep the measurements.\n\nBrayon Pieske is the founder of HexStellar and Trust Carbon Infrastructure, which operate between Brazil and the United States. He can be reached at [[email protected]](/cdn-cgi/l/email-protection) or on LinkedIn at https://www.linkedin.com/in/brayonpieske/.", "url": "https://wpnews.pro/news/hexstellar-founder-on-cutting-gpu-inference-energy-use-without-touching-the", "canonical_source": "https://startupfortune.com/hexstellar-founder-on-cutting-gpu-inference-energy-use-without-touching-the-model/", "published_at": "2026-08-12 08:03:04+00:00", "updated_at": "2026-08-12 08:14:26.270607+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-infrastructure", "ai-research"], "entities": ["HexStellar", "Brayon Pieske", "Trust Carbon Infrastructure", "StartupFortune"], "alternates": {"html": "https://wpnews.pro/news/hexstellar-founder-on-cutting-gpu-inference-energy-use-without-touching-the", "markdown": "https://wpnews.pro/news/hexstellar-founder-on-cutting-gpu-inference-energy-use-without-touching-the.md", "text": "https://wpnews.pro/news/hexstellar-founder-on-cutting-gpu-inference-energy-use-without-touching-the.txt", "jsonld": "https://wpnews.pro/news/hexstellar-founder-on-cutting-gpu-inference-energy-use-without-touching-the.jsonld"}}