Intelligence Raises $7.9M for Design Arena, Led by Index Ventures
Intelligence, the San Francisco company behind Design Arena, raised $7.9M in seed funding led by Index Ventures, with Conviction and A* participating. Intelligence identifies Y Combinator as another participant, while other coverage names Valkyrie instead. The 4th investor remains unconfirmed.
Founded in 2025 by Harvard classmates Grace Li, CEO, and Kamryn Ohly, CTO, Intelligence has built Design Arena around a deceptively simple problem: AI models can generate more things than humans could reasonably inspect, yet measuring whether those things are actually good remains considerably messier. Design Arena turns that mess into data. Users submit prompts, receive anonymized outputs from competing AI models and vote head-to-head. Intelligence reports approximately 5.3M–5.5M users across more than 190 countries, with ARR increasing from $5M to $60M in 6 months.
The funding matters because AI evaluation is moving beyond whether a model can produce an answer toward whether humans actually prefer the answer it produces. Intelligence is betting that taste, subjective as it sounds, can become measurable infrastructure.
What Happened
Index Ventures led Intelligence's $7.9M seed round, with Conviction and A* participating. Y Combinator, whose Summer 2025 batch included Intelligence, is identified by the company as another participant, although other coverage identifies Valkyrie instead. The 4th investor remains unconfirmed.
Grace Li and Kamryn Ohly founded Intelligence in 2025, weeks before graduating from Harvard University. Grace Li studied computer science and neuroscience, while Kamryn Ohly studied computer science and education. Both previously worked at Apple. Their original idea was not Design Arena. Grace Li and Kamryn Ohly were building an AI game-generation engine. The generated games functioned, but the outputs were not particularly appealing. That exposed a problem hiding behind the technical achievement: a machine successfully producing something does not mean a human wants it.
The founders built a this-or-that voting system to compare outputs. The diagnostic tool became the product. Startup mythology usually celebrates the original idea with suspiciously cinematic hindsight. Intelligence offers the more useful version. Sometimes the valuable company is hiding inside the instrument founders build to understand why their first idea is not working.
Why Design Arena Matters
Design Arena lets users submit prompts and compare anonymized outputs from competing AI models. Users vote head-to-head, allowing Intelligence to transform subjective human preferences into structured evaluation data. The anonymity matters. Claude, Grok, GLM, Gemini, Kimi and GPT-family models can compete without the user simply choosing the logo they already trust. Design Arena now covers more than 25 categories spanning websites, apps, games, images, video, audio, slides and 3D design.
This creates an unusual exchange. Users get access to an Arena where competing AI systems can be compared. Intelligence gets preference signals. AI labs can pay for private access to human-preference data that helps explain which outputs people actually favor. That is where Design Arena stops looking merely like a leaderboard.
Consumer participation can create proprietary preference data. Preference data can become useful to model developers. Model developers represent commercial demand. Intelligence sits between those groups and attempts to turn millions of tiny human judgments into infrastructure for AI evaluation. The crowd gets the Arena. The labs get signal. Intelligence gets a business.
From 50K Users to a Reported $60M ARR
Intelligence reports approximately 5.3M–5.5M users across more than 190 countries, compared with approximately 40K–50K users during Design Arena's first weeks in 2025. The company also reports ARR increasing from $5M to $60M in 6 months while operating with a team of 10. Frontier AI labs are paying for private access to human-preference data generated through evaluation.
Those numbers sharpen the strategic story behind the funding. Intelligence is not simply trying to attract people who enjoy comparing AI models. A conventional consumer product monetizes the user. Design Arena can derive commercial value from the judgment the user contributes. Each comparison potentially adds another signal about what people consider better design, better imagery, better applications or better creative output.
As generative models become more technically capable, evaluation becomes less binary. The question is no longer simply whether the model can generate a website, game or image. Increasingly, the question is which generated website, game or image a human would choose. Capability gets you onto the field. Preference starts deciding the score.
The Market for Human Judgment
AI benchmarking has historically rewarded measurable performance because measurable performance is convenient. Tests can determine whether software executes correctly, whether a model answers accurately or whether a system completes a defined task. Taste refuses to cooperate quite so neatly. A website can function perfectly and still look terrible. An image can satisfy a prompt and still feel wrong. A generated game can run without anybody wanting to play it. Creative AI exposes the distance between technical validity and human desirability.
Design Arena is attempting to quantify that distance. The company uses anonymized comparisons and tournament-style voting to produce rankings from human judgments. The underlying rankings use the Bradley-Terry model, a statistical approach to pairwise comparisons. That methodology gives Intelligence something more defensible than an endless stream of thumbs-up icons. The strategic asset is the structured preference dataset produced by repeated comparisons across models, categories and users.
Automated benchmarks can scale aggressively, but Intelligence argues that human evaluation captures subjective qualities automated systems struggle to measure and can reduce the influence of benchmarks that models may learn to optimize against. The irony is deliciously human: after spending fortunes teaching machines to generate creative work, the industry still needs people to tell the machines whether the work is any good.
Competitive Landscape
Investor interest suggests AI evaluation is becoming a category rather than a feature. LM Arena raised $150M in a Series A at a $1.7B valuation in January 2026. Yupp raised $33M and reached 1.3M users before shutting down in 2026. Those outcomes illustrate both sides of the opportunity. Human-preference platforms can attract users, capital and strategic attention, but scale does not automatically create a durable company.
For Intelligence, the defensibility question therefore extends beyond audience size. The harder question is whether Design Arena can produce preference data that AI labs consistently value enough to purchase, while maintaining a sufficiently broad and engaged human population to keep that data useful. That is a much more interesting business than winning the leaderboard popularity contest.
The $7.9M round gives Intelligence additional resources for model-evaluation infrastructure, expansion of its reviewer community and development of multimodal data pipelines. The company is also looking beyond visual design toward UI/UX, content-layout and writing-style preference evaluation, with multi-turn evaluation identified as a planned extension.
What This Signals for AI
The generative AI market has spent years obsessing over model capability. Intelligence represents another layer of the stack forming around model judgment. As competing systems become capable of producing acceptable code, images, applications, video and design, technical capability alone becomes less useful as a differentiator. Selection becomes harder precisely because generation becomes easier.
That creates demand for evaluation systems capable of answering increasingly subjective questions. Which interface feels better? Which generated image would somebody choose? Which application experience feels finished rather than merely functional? Those questions sound soft until companies start attaching revenue to the answers.
Intelligence's reported growth suggests human preference can become more than feedback. It can become a data product connecting consumer behavior with model development. For AI labs, better preference information can help determine what people actually value in generated output. For founders, Design Arena offers another lesson: the infrastructure required to evaluate a technology can become as strategically interesting as the technology being evaluated.
The Bigger Industry Shift
Grace Li and Kamryn Ohly started with an AI game engine and discovered that functioning software was not enough. Their attempt to understand why became Design Arena. That origin story is worth more than startup folklore because it mirrors the direction of generative AI itself. The first phase rewarded systems for making things possible. The next phase increasingly rewards systems for making things people prefer.
Intelligence is positioning itself inside that transition. Its advantage will not come simply from having millions of people click between AI outputs. The opportunity is converting those judgments into data that remains valuable as models improve, new modalities emerge and AI labs become increasingly sophisticated about evaluation.
Index Ventures, Conviction, A* and the other participants in the $7.9M seed are backing that possibility. AI keeps getting better at generating answers. Intelligence is building around the question that becomes more valuable with every new model: which answer do humans actually want?
Frequently Asked Questions
How much funding did Intelligence raise for Design Arena?
Intelligence raised $7.9M in seed funding for Design Arena. Index Ventures led the round, with Conviction and A* participating.
Who founded Intelligence and Design Arena?
Grace Li, CEO, and Kamryn Ohly, CTO, founded Intelligence in 2025. The Harvard graduates previously worked at Apple before building Design Arena.
What does Design Arena do?
Design Arena benchmarks generative AI models by showing users anonymized competing outputs and collecting head-to-head human-preference votes. Those judgments contribute to rankings and evaluation data.
How large is Design Arena?
Intelligence reports approximately 5.3M–5.5M users across more than 190 countries and more than 25 output categories. The company reports ARR increasing from $5M to $60M in 6 months with a team of 10.
How does Intelligence make money from Design Arena?
Intelligence says frontier AI labs pay for private access to human-preference data generated through evaluations on Design Arena.
How is Design Arena different from automated AI benchmarks?
Design Arena uses human judgments of anonymized AI outputs to evaluate subjective preferences, including design and creative quality, rather than relying exclusively on automated performance tests.
Who invested in Intelligence's $7.9M seed round?
Index Ventures led the $7.9M seed round, with Conviction and A* participating. Intelligence identifies Y Combinator as another participant, while other coverage identifies Valkyrie. The 4th investor remains unconfirmed.
Where the Money Moved
The intelligence briefing of the innovation economy. Funding, M&A, debt and fund closes, read as market signal rather than deal announcements.
Subscribe to Where the Money Moved








