Vals AI Raises $40M to Build AI's Evaluation Layer
Vals AI has raised a $40M Series A at a $400M valuation, giving the San Francisco AI-evaluation company fresh capital as enterprises demand better evidence about what models can actually do. a16z led the round, with existing investors 8VC, Pear VC, and Bloomberg Beta participating alongside new investors HRT Ventures and Next Ladder Ventures.
The financing matters because AI has an increasingly awkward measurement problem. Models keep improving on public leaderboards, but buyers still struggle to determine whether those systems can complete messy, multi-step work inside finance, law, healthcare, software development, and other high-stakes environments.
Vals AI is building the independent evaluation layer between model vendors and the organizations buying their technology. If the company succeeds, benchmarks become less like marketing scorecards and more like operating infrastructure for choosing models, tracking reliability, and deciding how much authority to give AI agents.
What Happened
Vals AI announced the Series A on August 13, 2026. The official announcement identified a16z as the lead investor, with 8VC, Pear VC, Bloomberg Beta, HRT Ventures, and Next Ladder Ventures completing the disclosed syndicate.
The company was founded by Rayan Krishnan, CEO, and Langston Nashold, CTO, who studied computer science at Stanford. Vals previously disclosed a $5M seed round, but its latest announcement did not state an official lifetime-funding total, so the cleanest comparison is the new $40M round against that prior financing.
Vals paired the funding announcement with three releases. Vals Smith became generally available for building custom coding benchmarks from GitHub repositories, the company introduced frontier-risk evaluations, and Vals 2.0 expanded its public benchmark and model-comparison surface.
Why AI Evaluation Is Becoming Infrastructure
The old benchmark model worked when AI systems improved slowly enough for a stable exam to remain useful. That bargain is breaking as frontier models saturate public datasets, train on material that may include test questions, and get optimized against the same leaderboards used to declare progress.
Enterprise buyers face a different problem from model labs. They do not need to know whether a system can win a familiar academic contest; they need to know whether it can analyze a credit agreement, complete legal research, modify a real codebase, or operate across a long business process without producing expensive surprises.
That gap grows as AI moves from short responses to agentic work lasting hours or days. A weak model choice can waste tokens and employee time, but it can also introduce flawed analysis, unreliable automation, and customer-facing errors that are much harder to dismiss than a bad chatbot answer.
How Vals AI Tests Real-World Performance
According to the company’s evaluation methodology, Vals works with domain experts to translate professional workflows into benchmarks for finance, healthcare, law, coding, and other fields. It keeps key test sets private to reduce contamination and gaming, while measuring accuracy alongside cost, latency, failure patterns, and qualitative performance.
The company also treats benchmarks as perishable. When a test stops separating stronger models from weaker ones, Vals can retire it and introduce a harder evaluation, an approach designed to keep measurement aligned with the moving frontier rather than preserve a leaderboard after its signal has faded.
Vals Smith applies that idea to software teams. It identifies merged pull requests in a public or private GitHub repository, converts suitable changes into self-contained tasks with hidden tests and reproducible setups, and then measures whether coding models or agents can complete the work without access to the original fix.
What the Traction Says
Vals reported that its revenue grew 8x compared with all of 2025, its customer base doubled, and its team tripled in six months. Those figures are company-reported rather than independently audited, but they suggest that model evaluation is moving from research curiosity to a budgeted enterprise need.
The company also says its results have been cited in model cards from OpenAI, Anthropic, Google, Meta, and xAI. That matters because a credible evaluation platform needs influence on both sides of the market: model builders must respect its measurements, while enterprise buyers must trust those results enough to guide deployment decisions.
The company’s work with the U.S. Department of Commerce and members of Congress adds another dimension. Governments need ways to assess frontier capabilities and cyber risk without relying exclusively on the claims of the companies building the systems under review.
The Market Signal Behind a $400M Valuation
The $400M valuation is a wager that measurement becomes more valuable as model capability becomes more abundant. When vendors know more than buyers and have every incentive to present favorable results, markets tend to create demand for an independent institution that can make performance legible.
That does not make Vals a guaranteed winner, and the evaluation market will not belong to one methodology. Public leaderboards, crowdsourced preference testing, private enterprise evaluations, safety benchmarks, and task-specific coding tests all answer different questions, while model labs will keep improving their own internal measurement systems.
Vals is making a sharper bet: real-world, adaptive, private testing can become the trusted layer that helps organizations choose among those systems. The Series A gives the company resources and investor backing to extend that thesis while its new products put the idea in front of developers, enterprises, model labs, and policymakers.
What Operators Should Watch Next
The first test is whether custom evaluations change buying behavior. If enterprises use repository-specific and workflow-specific benchmarks to select models, negotiate contracts, or set deployment controls, evaluation will move closer to procurement infrastructure and further away from leaderboard entertainment.
The second test is independence. Vals must keep winning trust from model developers whose systems it grades and from buyers who need results that are current, reproducible, and resistant to gaming. That balancing act gets harder as the company grows and its measurements carry more commercial weight.
The larger shift is already visible. AI models are getting cheaper, faster, and more capable, but evidence about reliability remains scarce. Vals AI has raised $40M on the belief that the next phase of the market will reward the companies that can prove what intelligence is worth, not only the companies producing more of it.
AI & Machine Learning funding, last 30 days
DevCuration's funding database tracked 14 AI & Machine Learning rounds totaling $1.8B in disclosed capital over the past 30 days. Recent deals we covered:
- Etched Raises $700M at $21B as Inference Race Accelerates$700M · Aug 20
- Wispr Raises $280M Series B at $2B ValuationSeries B · $280M · Aug 18
- Worldscape Raises $10M Seed Extension for Mission AISeed Extension · $10M · Aug 18
- Palona AI Discloses $28.315M Equity Financing$28.315M · Aug 18
- Higgsfield Raises $400M Series B at $5.4B Valuation$400M · Aug 18
Frequently Asked Questions
What does Vals AI do?
Vals AI builds independent benchmarks and evaluation infrastructure for AI models and agents. Its tests focus on realistic work in areas such as finance, law, healthcare, and software development, using private test sets to reduce contamination and gaming.
Who invested in Vals AI's Series A?
a16z led the $40M Series A. Existing investors 8VC, Pear VC, and Bloomberg Beta participated alongside new investors HRT Ventures and Next Ladder Ventures.
Why are private AI benchmarks important?
Private benchmarks make it harder for test material to leak into training data or become a target for narrow optimization. That can give enterprises a clearer view of how models perform on unfamiliar, real-world tasks.
What is Vals Smith?
Vals Smith creates coding benchmarks from merged pull requests in public or private GitHub repositories. It turns suitable changes into reproducible tasks with hidden tests so teams can compare models and coding agents against work drawn from their own codebase.
Where the Money Moved
The intelligence briefing of the innovation economy. Funding, M&A, debt and fund closes, read as market signal rather than deal announcements.
Subscribe to Where the Money Moved