Wafer Raises $40M to Automate AI Inference Optimization
Every new AI accelerator adds theoretical choice and practical work. Wafer raised a $40M Series A to automate the performance engineering that determines whether a model, serving engine, kernel stack, and chip can deliver reliable production economics. Marathon and Chemistry co-led the financing, with Wing Venture Capital, AMD Ventures, and Outset Capital participating alongside existing investors Fifty Years and Y Combinator.
The round matters because the hardware beneath AI is getting more diverse while the work required to use it well remains deeply specialized. Wafer is building agents that learn how a production workload behaves, search across the serving stack, and keep recalculating the best deployment as traffic, models, and accelerators change.
The September 1, 2026 announcement follows a $4M Seed disclosed in April, bringing Wafer’s total disclosed funding to $44M. The new capital will help the San Francisco company automate more of the inference-optimization loop and expand the team doing that work.
What Wafer raised and who backed it
Wafer announced the $40M Series A on September 1, 2026. Marathon Management Partners and Chemistry co-led the round, while Wing Venture Capital, AMD Ventures, and Outset Capital joined as participants. Fifty Years and Y Combinator, both existing investors, also invested again.
The financing arrived less than 5 months after Wafer announced its $4M Seed round, which was led by Fifty Years with participation from Liquid2 and Y Combinator. Primary sources did not disclose a valuation for the Series A. Wilson Sonsini, which advised Wafer on the transaction, independently confirmed the amount, round, investor roster, and September 1 announcement date in its transaction note.
Why inference optimization is becoming a software layer
AI inference is the work that happens when a trained model serves a live request. The invoice may name an accelerator, but the economics depend on far more than the chip: model architecture, memory, batching, caching, scheduling, latency targets, traffic shape, and the software connecting those pieces all change the cost and performance of the endpoint.
That complexity has traditionally created demand for performance engineers who can profile workloads, tune kernels, test serving configurations, and keep the system stable. The work is expensive, specialized, and often repeated when a model changes or new hardware becomes available. A benchmark can show what happened under one set of conditions; a production team still has to make those gains survive customer traffic.
The accelerator market adds pressure. NVIDIA’s CUDA ecosystem remains the default for much of the industry, while AMD GPUs, Google TPUs, AWS Trainium, Cerebras systems, and custom silicon expand the menu. More choice can improve economics, but only when software makes the hardware usable for a specific workload. Otherwise, every new option arrives with its own optimization tax.
What Wafer is building
Wafer describes its product as continual inference. The system learns a workload’s traffic patterns and performance constraints, then searches across the model, inference engine, kernels, and hardware for a stronger deployment. The search can include batching, caching, quantization, speculative decoding, memory management, scheduling, and chip selection.
The important shift is temporal as much as technical. Most optimization happens before deployment or through a services-heavy engagement, then begins aging as traffic changes. Wafer wants the optimization loop to remain active, measuring the workload again and deploying only changes that preserve correctness and reliability requirements.
That makes the product less like another generic inference endpoint and more like a control layer for deployment economics. The customer is buying the ability to use a broader hardware market without assembling a specialist performance team around every model. The business case depends on Wafer proving that those decisions remain trustworthy when they move from a benchmark into live production.
Early traction and the evidence boundary
Wing Venture Capital reports that Wafer reached approximately $8M in annualized revenue in roughly 3 months. Wing also says Wafer’s agents can tune more than 100 serving parameters and compress optimization work that may take days or weeks into hours. These figures come from a participating investor and should be read as investor-reported evidence, not an independent audit.
Wafer and Wing also publish workload-specific performance comparisons across AMD and NVIDIA hardware. Those tests illustrate why model architecture, memory capacity, topology, software, and rental prices can change the economic result. They do not prove that one accelerator wins every workload, and Wafer’s larger opportunity depends on avoiding that kind of universal claim.
The strategic participation of AMD Ventures sharpens the point. Wafer can benefit from the growth of non-NVIDIA infrastructure while still needing credibility across a heterogeneous market. If customers see the system as a neutral optimizer rather than an advocate for one silicon ecosystem, the company can compete for the decision layer that routes demand among them.
The founders behind Wafer
Co-founders Emilio Andere and Steven Arellano met at the University of Chicago, where they were classmates and later roommates. Current Y Combinator records identify Andere as co-founder and CEO and Arellano as co-founder. The company participated in YC’s Summer 2025 batch and lists San Francisco as its headquarters.
Andere studied mathematics, researched weather models at Argonne National Laboratory, and published machine-learning security research. Arellano’s background includes high-performance computing and AI infrastructure work at Google, Two Sigma, and Sei Labs. Their experience fits the problem Wafer has chosen: the company has to understand mathematical performance, low-level systems behavior, and the commercial consequences of both.
What the $40M changes
Wafer says the Series A will help automate more of the inference-optimization loop and advance infrastructure that keeps improving after deployment. Hiring evidence on the company’s live channels points to expansion across technical staff, growth, operations, and go-to-market work, which turns the financing into both a research agenda and a company-building exercise.
The market signal reaches beyond Wafer. As model providers and application companies consume more inference, chip diversity alone will not determine where that spending lands. Control layers that can translate real traffic into reliable performance-per-dollar decisions may become the brokers between models and silicon.
Wafer now has capital, early investor-reported traction, and a syndicate spanning venture firms and a major chip company. The part still moving is the production record: each new workload gives its agents another opportunity to prove that better inference economics can be discovered continuously, even after the hardware has arrived and the endpoint is already carrying customer expectations.
AI Infrastructure funding, last 30 days
DevCuration's funding database tracked 26 AI Infrastructure rounds totaling $7B in disclosed capital over the past 30 days. Recent deals we covered:
- Physical Superintelligence Raises $58M for AI PhysicsSeed · $58M · Sep 2
- Visko Raises $10M for Real-Time AI Model OrbisPre-Seed · $10M · Sep 1
- TrustedRouter Raises $1.25M for Verifiable AI RoutingSeed · $1.25M · Sep 1
- a16z Raises $1.1B Machine Age Fund for Physical AI$1.1B · Aug 31
- Lambda Closes $926M Loan for AI Cloud InfrastructureTerm Loan B · $926M · Aug 28
Frequently Asked Questions
What problem does Wafer solve for AI infrastructure teams?
Wafer automates performance engineering across models, serving engines, kernels, and hardware. The goal is to keep improving throughput, latency, reliability, and cost per token as a production workload changes.
Why does a multi-silicon inference market need an optimization layer?
Different workloads can perform differently across accelerators, memory systems, serving engines, and traffic patterns. An optimization layer can help teams use that wider hardware market without manually retuning every deployment.
What evidence supports Wafer’s early traction?
Wing Venture Capital reports that Wafer reached approximately $8M in annualized revenue in roughly 3 months and that its agents tune more than 100 serving parameters. Those figures are investor-reported rather than independently audited.
What will Wafer use the Series A funding for?
Wafer says the capital will help automate more of the inference-optimization loop and advance infrastructure that continually improves. Current company channels also show hiring across technical, growth, operations, and go-to-market roles.
What should AI operators watch next?
The important evidence will come from production workloads: whether Wafer can preserve correctness and reliability while improving economics across changing models, traffic patterns, and hardware environments.
Where the Money Moved
The intelligence briefing of the innovation economy. Funding, M&A, debt and fund closes, read as market signal rather than deal announcements.
Subscribe to Where the Money Moved