Infinity Raises $15 Million Seed to Automate AI Inference Software for Any Chip Infinity, a San Francisco-based AI infrastructure startup, announced a $15 million seed round at a $100 million post-money valuation to scale its platform that automates the adaptation of inference software for any AI chip. The round was led by Touring Capital and Principal VC, with participation from executives at chip companies and researchers from OpenAI and Anthropic. Infinity's core offering, Ignition, is an autonomous AI research agent that generates and optimizes low-level compute kernels, aiming to make any chip inference-ready quickly. San Francisco-based AI infrastructure startup Infinity https://www.streetinsider.com/Business%2BWire/Infinity%2BRaises%2B%2415%2BMillion%2Bin%2BSeed%2BFunding%2Bto%2BBuild%2Bthe%2BSoftware%2BLayer%2BThat%2BMakes%2BAny%2BAI%2BChip%2BInference-Ready/26788462.html?utm source=openai announced today it has raised a $15 million seed round at a $100 million post-money valuation. The funding will scale its platform for automating the adaptation of inference software for new AI chips. The round was backed by Touring Capital https://touringcapital.com/ and Principal VC https://www.principal.vc/ , with additional participation from executives at chip companies, researchers from OpenAI and Anthropic, and other angel investors, according to the company. Samir Kumar, General Partner at Touring Capital, and Songyee Yoon, founder and Managing Partner at Principal Venture Partners, are quoted in the announcement regarding their investment. Jeremy Nixon https://jeremynixon.github.io/ , Infinity's founder and CEO, previously served as a research software engineer at Google Brain and founded the AGI House hacker network. He also founded Omniscience, a startup reported by Forbes in 2021, and holds multiple ML research publications, including co-authorships at ICML 2019 and an ICLR 2026 workshop item. Nixon, who studied at Harvard University, launched Infinity in August 2025 to address a bottleneck in AI chip adoption. Infinity's core offering, named "Ignition," is described as an autonomous AI research agent and automated research platform. This system generates, tests, debugs, and optimizes low-level compute kernels--the inference software crucial for performance on new AI chips. The company positions Ignition as a software layer designed to write inference stacks across diverse hardware architectures, aiming to quickly make any chip inference-ready. Instead of upfront licensing, Infinity plans a performance-based commercial model, sharing gains with partners. Traction and Partnerships Infinity claims to have achieved "millions of dollars in ARR" from its chip design partnerships, though this figure is self-reported by the company and lacks third-party verification. Its work with chip partner d-Matrix https://d-matrix.ai/ provides an example of its capabilities. According to Infinity, its agents achieved up to 92% of a new d-Matrix chip's theoretical peak performance on tensor-parallel matrix multiplications within 10 hours of initial hardware access. The company further states it had three frontier models--Qwen3, Qwen3.5, and Gemma4--running end-to-end on d-Matrix hardware within 10 days. For the Qwen3-8B model, Infinity claims a 34% improvement in tokens per second within one day, outperforming established solutions like vLLM. These performance benchmarks are derived from company materials and have not been independently verified. As of the announcement, Infinity reportedly employs approximately 26 people, according to TechCrunch. Competitive Landscape Infinity enters a market dominated by NVIDIA's CUDA ecosystem https://developer.nvidia.com/cuda-zone , which has long been the de facto standard for AI hardware and software integration. Infinity aims to provide an alternative or complementary path for chipmakers, directly challenging NVIDIA's software advantage. Other projects and companies operating in related spaces include open-source compiler/runtime frameworks like Apache TVM https://tvm.apache.org/ and commercial entities like OctoML https://octoml.ai/ that also focus on optimizing models across various hardware. Intel's oneAPI/Level Zero and AMD's ROCm offer their own cross-hardware deployment tools. While inference platforms such as vLLM https://github.com/vllm-project/vllm and Triton Inference Server https://github.com/triton-inference-server/server manage inference at a higher level, Infinity's focus on automated kernel generation and RTL/ISA-agnostic kernel writing carves out a specific niche.