{"slug": "exclusive-iterate-ais-lifeboat-runs-up-to-six-times-more-ai-agent-sessions-per", "title": "Iterate.ai’s Lifeboat runs up to six times more AI agent sessions per GPU", "summary": "Iterate Studio Inc. launched Lifeboat, an LLM inference engine with built-in confidential computing that the company says fits two to six times as many concurrent AI agent sessions per GPU. In Iterate.ai's own testing on a single Nvidia RTX PRO 6000 Blackwell GPU running a Qwen 30B-A3B model, Lifeboat held 2,048 concurrent sessions with every request completing, versus half that with its optimizations off, and reached 8,714 tokens per second against 4,965. Lifeboat is generally available now, with a free Developer License for up to two inference servers on one node, a Standard License at $49.99 per month, and a Confidential Computing edition at $499.99 per month that withholds service until hardware attestation passes.", "body_md": "### Exclusive: Iterate.ai’s Lifeboat runs up to six times more AI agent sessions per GPU\n\nEnterprise artificial intelligence software company [Iterate Studio Inc.](https://iterate.ai/) today launched Lifeboat, an inference engine for large language models that has confidential computing built in.\n\nIterate.ai says the software fits two to six times as many concurrent AI agent sessions on each graphics processing unit.\n\nLifeboat is seeking to take on a memory problem that agents create at scale. A chatbot question usually triggers one model call. An agent working through a single task can make dozens, and its context window grows with each one, so hundreds of agents on shared hardware fill the key-value cache quickly. Iterate.ai says standard inference engines can stall at four or five long-context requests running at once.\n\nTeams that hit that limit tend to buy more GPUs or move the work to per-token cloud services and hand their data to a third party. Brian Sathianathan, Iterate.ai’s co-founder and chief technology officer, said banks, insurers and health systems want to run agents on their own data “inside their own walls” without doubling GPU spending.\n\nLifeboat’s approach starts with fair scheduling and admission control, which give each session its share of the GPU so a heavy document-processing agent cannot crowd out interactive ones. Optimizations to the key-value cache double its effective capacity while model weights stay at full precision, according to the company.\n\nOn mixture-of-experts models, only the experts a given agent needs get loaded. Every session also runs inside its own security capsule with filtering, token budgets and sandboxed execution.\n\nIterate.ai’s own testing used a single Nvidia Corp. RTX PRO 6000 Blackwell GPU running a Qwen 30B-A3B model, and Lifeboat held 2,048 concurrent sessions on it with every request completing. The same engine with Lifeboat’s optimizations switched off topped out at half that number and managed 4,965 tokens per second against 8,714 with them on. The widest gap showed up in a memory-pressure run of 128 sessions sending 18,000-token requests, in which Lifeboat kept 99th-percentile time to first token at 1.5 seconds. The baseline needed 189.\n\n“Before any enterprise buys more GPUs for its agents, it should find out what the ones it already owns can do,” said Jon Nordmark, co-founder and chief executive of Iterate.ai. “A data center or neo-cloud that doubles concurrent sessions per card gets that capacity back without adding racks or power.”\n\nLifeboat’s top-tier Confidential Computing edition will not serve a request until hardware attestation passes. That check covers trusted execution features in Advanced Micro Devices Inc. and Intel Corp. processors, plus Nvidia’s confidential computing mode on the H100, B200, GB300 and other supported GPUs. Model weights are sealed inside the trusted execution environment and stay encrypted in use, whether the engine sits in a cloud confidential virtual machine or on customer-owned hardware.\n\nEarly access to Lifeboat drew [thousands of downloads](https://github.com/IterateAI/lifeboat-releases/releases/tag/engine-v2.2.54) within days, the company said, and the engine is [generally available now](https://iterate.ai/lifeboat). A free Developer License covers noncommercial and evaluation use on up to two inference servers on a single node. Customers who want professional support can move to the Standard License at $49.99 per month. The attestation features sit in the Confidential Computing edition, which costs $499.99 per month and, like the Standard tier, can be tried free for seven days without a credit card.\n\nNordmark and Sathianathan discussed the company’s Generate platform and its partnership with NetApp Inc. on theCUBE, SiliconANGLE Media’s livestreaming studio, at NetApp’s Insight conference [last week](https://siliconangle.com/2026/10/02/open-weight-models-power-private-ai-netapp-iterate-netappinsight/):\n\n##### Image: Iterate.ai\n\n# A message from John Furrier, co-founder of SiliconANGLE:\n\nSupport our mission to keep content open and free by engaging with theCUBE community. **Join theCUBE’s Alumni Trust Network**, where technology leaders connect, share intelligence and create opportunities.\n\n- **15M+ viewers of theCUBE videos** , powering conversations across AI, cloud, cybersecurity and more\n- **11.4k+ theCUBE alumni** — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network\n\n### Are you an AWS customer?  Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: [https://siliconangle.com/aws-marketplace/](https://siliconangle.com/aws-marketplace/)\n\n##### **About SiliconANGLE Media**\n\n[SiliconANGLE](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fsiliconangle.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=SiliconANGLE&index=9&md5=646b1b564e2259100a2b8638aab0a552),\n\n[theCUBE Network](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecube.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Network&index=10&md5=7de2a85f95ab4a4a495cede20b8cb1da),\n\n[theCUBE Research](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fthecuberesearch.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Research&index=11&md5=7bb33676722925eb57d588ec343e4f6f),\n\n[CUBE365](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.cube365.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=CUBE365&index=12&md5=d310fb35919714e66ad8d42c9c0c1bc6),\n\n[theCUBE AI](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecubeai.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+AI&index=13&md5=b8b98472f8071b23ebb10ab9a8dd0683)and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.\n\nFounded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.", "url": "https://wpnews.pro/news/exclusive-iterate-ais-lifeboat-runs-up-to-six-times-more-ai-agent-sessions-per", "canonical_source": "https://siliconangle.com/2026/10/05/exclusive-iterate-ais-lifeboat-runs-up-to-six-times-more-ai-agent-sessions-per-gpu/", "published_at": "2026-10-05 13:00:47+00:00", "updated_at": "2026-10-05 13:18:31.328478+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "ai-agents", "ai-products", "ai-safety"], "entities": ["Iterate Studio Inc.", "Lifeboat", "Brian Sathianathan", "Jon Nordmark", "Nvidia Corp.", "Qwen 30B-A3B", "Advanced Micro Devices Inc.", "Intel Corp."], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/exclusive-iterate-ais-lifeboat-runs-up-to-six-times-more-ai-agent-sessions-per", "markdown": "https://wpnews.pro/news/exclusive-iterate-ais-lifeboat-runs-up-to-six-times-more-ai-agent-sessions-per.md", "text": "https://wpnews.pro/news/exclusive-iterate-ais-lifeboat-runs-up-to-six-times-more-ai-agent-sessions-per.txt", "jsonld": "https://wpnews.pro/news/exclusive-iterate-ais-lifeboat-runs-up-to-six-times-more-ai-agent-sessions-per.jsonld"}}