Infinity Raises $15M Seed to Build AI Inference Software for Next-Generation Chips
Infinity, an AI infrastructure startup building inference software for next-generation AI chips, has raised a $15M seed round at a reported $100M valuation. The San Francisco-based company is developing Ignition, an autonomous AI research agent that generates, tests, and optimizes inference code for new silicon so hardware teams can move from first access to production-ready performance much faster.
The round includes backing from Touring Capital and Principal Venture Partners, alongside researchers from OpenAI and Anthropic. Founded by Jeremy Nixon, a former Google Brain researcher and AGI House co-founder, Infinity is positioning itself as the software layer for a hardware market that continues to become both more ambitious and more fragmented.
The announcement matters because inference is becoming the daily operating workload of enterprise AI. Training captures the headlines, but inference is where models meet users, margins, latency, and infrastructure costs every single day.
What Happened
AI chips do not become commercially useful simply because the silicon is impressive. They require mature software, optimized kernels, model support, compiler tooling, and deployment infrastructure before customers trust them in production. That software layer has become one of the biggest constraints facing emerging hardware vendors.
According to Infinity's funding announcement, the company raised $15M in seed funding to build the software layer that makes any AI chip inference-ready. Infinity says Ignition can engineer custom inference software for new silicon in days rather than years, dramatically shortening the time required to evaluate, optimize, and deploy new AI accelerators.
Founded in August 2025, Infinity remains an early-stage company, but its thesis is highly specific. Rather than promoting a broad AI infrastructure narrative, it aims to automate the specialized systems engineering required to make frontier AI models run efficiently across new chip architectures, compressing work that traditionally depends on scarce engineering expertise.
Why AI Inference Software Has Become the Real Battleground
Hardware captures attention because it is tangible. Software determines whether that hardware can actually win production workloads, and the AI market has already demonstrated how durable that advantage becomes when developers build around a mature software ecosystem.
For more than two decades, NVIDIA benefited from more than increasingly powerful chips. Its software ecosystem gave developers confidence that models, libraries, and production workloads would run reliably, creating a competitive advantage that became just as valuable as the hardware itself.
Infinity is approaching that challenge from a different direction. Instead of manually engineering inference libraries for every new processor, Ignition automatically generates, tests, and optimizes low-level compute kernels tailored for emerging chip architectures, with the goal of reducing development timelines from months or years to days.
That distinction matters because inference is expected to account for the majority of AI compute spending as enterprise adoption expands. Training foundation models remains expensive and highly visible, but inference becomes the recurring operational workload businesses pay for every time a model serves a customer, executes a workflow, summarizes information, or responds inside a production application.
Performance Matters More Than Promises
Infrastructure startups often sell vision before evidence. Infinity distinguishes itself by publishing measurable technical results, particularly through work with d-Matrix and its Corsair accelerator.
According to the company, Ignition achieved up to 92% of the Corsair accelerator's theoretical peak performance within 10 hours of first hardware access. Infinity also reports bringing frontier models, including Qwen3, Qwen3.5, and Gemma4, to production-quality inference on the platform within 10 days using software generated from scratch.
The company further reports that its inference engine outperformed vLLM by more than 34% in tokens per second on Qwen3-8B under identical parameters. It also says automated optimization increased throughput from roughly 1,400 tokens per second to more than 20,000 tokens per second.
Those benchmarks naturally warrant independent validation, particularly because benchmarking methodology can significantly influence performance comparisons. Even so, the broader signal is clear: Infinity is attempting to automate systems engineering work that has traditionally required specialists across compilers, AI runtimes, hardware architecture, and low-level performance optimization.
Why Investors Are Paying Attention
Infrastructure investing rarely rewards incremental workflow improvements. Investors typically look for technologies capable of changing how an entire market operates, and Infinity's thesis sits squarely at the intersection of semiconductor innovation and enterprise AI deployment.
Touring Capital and Principal Venture Partners are backing a company built around a practical industry challenge: more AI hardware options are emerging, but software maturity still determines whether those chips can compete for production workloads. That makes Infinity relevant not only to chip startups, but also to model developers, cloud providers, enterprise AI teams, and infrastructure organizations seeking greater hardware flexibility.
The participation of researchers from OpenAI and Anthropic also reflects growing interest in infrastructure that expands hardware choice instead of concentrating the market around a single software ecosystem. Nixon's experience at Google Brain and AGI House further reinforces that narrative by connecting frontier AI research with the deployment challenges that emerge when advanced models reach production.
What This Signals for AI Infrastructure
The AI industry is gradually moving beyond a conversation centered solely on larger models and faster chips. The next phase is increasingly defined by operational efficiency, deployment speed, hardware flexibility, and the software infrastructure that makes advanced hardware practical at scale.
Infinity reflects that evolution. Rather than competing directly with chip manufacturers, the company aims to become an enabling software layer that helps more hardware platforms reach production readiness faster while giving customers greater freedom to evaluate hardware on performance instead of software maturity.
That shift matters because AI infrastructure rarely remains static. Hardware innovation continues accelerating across startups and established semiconductor companies alike, and businesses that reduce integration complexity, automate optimization, and shorten deployment timelines may influence adoption just as much as the hardware vendors themselves.
There is also a broader lesson for founders raising capital today. Investors continue backing ambitious technical visions, but increasingly expect measurable engineering evidence alongside the narrative. Infinity's announcement leads with technical outcomes rather than generic AI enthusiasm.
Infinity's $15M seed round is more than another AI funding announcement. It reflects growing confidence that the next competitive frontier will belong not only to companies building faster chips, but also to those building the software that allows every promising chip to perform like it was ready for production from day one.
AI Infrastructure funding, last 30 days
DevCuration's funding database tracked 37 AI Infrastructure rounds totaling $31.7B in disclosed capital over the past 30 days. Recent deals we covered:
- Gritt Raises $26M to Bring Physical AI to ConstructionSeries A · $26M · Jul 22
- MyDecisive Raises $12M Seed to Cut Enterprise Observability Costs in the AI EraSeed · $12M · Jul 22
- Runta Raises $20M Seed Led by Andreessen Horowitz for AI Agent Runtime SecuritySeed · $20M · Jul 21
- General Compute Secures $400M Upper90 Debt Facility for AI InferenceDebt · $400M · Jul 21
- Nebius Secures $775M Debt Facility to Expand AI Cloud InfrastructureDebt · $775M · Jul 20
Frequently Asked Questions
What does Infinity do in AI infrastructure?
Infinity builds AI inference software for new silicon. Its autonomous AI research agent, Ignition, generates, tests, and optimizes low-level inference code so new AI chips can reach production-ready performance faster.
Why does Infinity's $15M seed round matter?
The round signals investor demand for software that can make emerging AI hardware easier to deploy. As inference becomes a larger share of AI compute usage, the software layer that improves performance and reduces deployment time becomes strategically important.
What is Ignition?
Ignition is Infinity's autonomous AI research agent for generating inference software on new chip architectures. Infinity says it can compress work that often takes months or years into days while producing optimized kernels and runtime support.
Who backed Infinity's seed round?
The round included Touring Capital and Principal Venture Partners, along with researchers from OpenAI and Anthropic. The company reported a $15M seed round at a $100M valuation.
Why is AI inference software important?
Inference is the recurring workload that runs AI models in real products, not just during model training. Better inference software can reduce friction between new hardware and production deployment, which may broaden competition across the AI infrastructure market.









