Modulate Raises $25M for Audio-Native Voice AI
Modulate has raised $25M in new funding to expand its audio-native artificial-intelligence platform for developers and enterprises. Future Ventures led the September 28, 2026 financing, with returning investors Hyperplane and Lakestar participating. Modulate says the round brings its total funding to $60M, but the company did not disclose a round series or valuation. The Greater Boston company is building around a simple technical and commercial problem: a transcript preserves words while discarding much of the evidence carried by sound. Tone, emotion, timing, emphasis, overlap, background audio, speaker behavior, and synthetic-voice signals can all matter when software is deciding whether a customer is frustrated, a caller is attempting fraud, or an AI agent is behaving as intended.
That gap is becoming more expensive as voice agents move into live customer interactions and synthetic voices become easier to produce. Modulate is using the new capital to expand research, engineering, developer access, and partnerships around a shared audio-intelligence layer rather than asking every voice product team to build those capabilities from scratch.
What Modulate Raised
Modulate's official announcement names Future Ventures as the lead investor and Hyperplane and Lakestar as participants. The financing follows a $30M Series A led by Lakestar in 2022. Modulate reports $60M in cumulative funding, although the new announcement does not provide a complete round-by-round history, ownership terms, valuation, or board changes.
The absence of a round label matters. Secondary databases may classify the financing as a Series B, but the primary company sources describe it only as new funding. The accurate public record is therefore a $25M venture round with an undisclosed series. Modulate plans to invest across AI and machine-learning research, product and engineering, developer relations, APIs, SDKs, deployment options, and partner integrations. The goal is to make sophisticated audio understanding available wherever developers are building voice products.
From Game Chat to Audio Infrastructure
Carter Huffman and Mike Pappas founded Modulate in 2017 after meeting at MIT. According to the company's current leadership page, Huffman is CEO and Pappas is chairman. That is a recent change from earlier 2026 materials that listed Pappas as CEO and Huffman as CTO.
Modulate's first large production problem was online game voice chat. Its ToxMod system was designed to detect harassment, threats, grooming, and other harmful behavior while preserving the context that separates abuse from jokes or ordinary competitive banter. In 2023, Modulate announced an Activision partnership to support real-time voice moderation in Call of Duty.
Gaming gave Modulate a difficult acoustic environment in which to learn. Real conversations include interruptions, laughter, poor microphones, multiple languages, music, background noise, and meaning that depends on how something was said. Those conditions resemble the problems now appearing in contact centers, fraud systems, voice agents, and identity workflows more than a clean speech-to-text demo does.
How Velma Listens Beyond the Transcript
Velma is Modulate's broader voice-intelligence platform. The company describes its Ensemble Listening Model architecture as a system of more than 100 specialized audio models that combine different signals into higher-level judgments about a conversation.
The current developer documentation includes APIs for conversation analysis, transcription, deepfake detection, emotion detection, accent and language detection, audio-event detection, AI-music detection, and PII or PHI redaction. Some products work on complete recordings, while others support streaming audio for live decisions.
This is a different product thesis from running a transcript through a general-purpose language model. Transcription remains useful and is part of Modulate's stack, but it becomes one signal among many. The practical question for a buyer is whether the missing acoustic context changes the decision the application has to make.
The Evidence Behind the Round
Modulate says its models now analyze more than 10M hours of audio each month and have processed more than 600M hours in total. The company also reports leading results on public Hugging Face transcription and deepfake-detection benchmarks, including a 98.9% result for its synthetic-voice detector.
Those figures should be read with appropriate boundaries. Public benchmarks provide more evidence than a private demo, and years of game-moderation deployment show that the technology has operated in noisy, adversarial environments. They do not establish accuracy across every language, customer population, fraud tactic, call-center system, or voice-agent failure mode. Revenue, customer count, retention, and audited enterprise outcomes were not disclosed.
The policy environment reinforces that broader proof burden. The Federal Trade Commission has described voice-cloning harms as a problem requiring prevention, authentication, real-time detection, and post-use evaluation rather than one technical fix. The Federal Communications Commission has applied existing artificial-voice restrictions to AI-generated robocalls. Detection technology will sit inside a larger operating system of consent, identity, escalation, human review, and enforcement.
What the $25M Has to Prove
Modulate's opportunity is to become infrastructure beneath the growing voice economy. Developers building assistants, contact-center tools, fraud systems, safety products, and media applications may prefer one audio-intelligence layer to a patchwork of specialized models and internal research projects.
The challenge is that the most valuable use cases are also the least forgiving. A false positive can interrupt a legitimate customer, burden an agent, or trigger an unnecessary investigation. A missed synthetic voice or escalation signal can carry financial, safety, and reputational consequences. Buyers will need evidence that the system remains accurate, affordable, explainable, and operationally useful in their own audio.
The round gives Modulate more capacity to turn nearly a decade of listening research into a developer platform. Its next measure of progress will be the quality of the decisions customers can make with signals that disappeared when voice was treated as text alone.
Frequently Asked Questions
What did Modulate announce in September 2026?
Modulate announced $25M in new funding on September 28, 2026. Future Ventures led the financing, with Hyperplane and Lakestar participating, and Modulate says the round brings total funding to $60M.
Who invested in Modulate's $25M round?
Future Ventures led the round. Returning investors Hyperplane and Lakestar also participated.
What does Modulate's voice AI platform do?
Modulate's Velma platform analyzes raw audio signals as well as spoken words. Its APIs cover capabilities such as transcription, conversation analysis, deepfake detection, emotion and accent detection, audio-event detection, and sensitive-data redaction.
How will Modulate use the new funding?
Modulate says it will expand AI and machine-learning research, product and engineering, developer relations, APIs, SDKs, deployment options, partner integrations, and its broader developer ecosystem.
Why does audio-native analysis matter for voice AI?
A transcript preserves words but may lose tone, timing, emotion, identity, background audio, and synthetic-voice signals. Those details can affect fraud detection, voice-agent supervision, customer experience, moderation, and other decisions made from live conversations.
Where the Money Moved
The intelligence briefing of the innovation economy. Funding, M&A, debt and fund closes, read as market signal rather than deal announcements.
Subscribe to Where the Money Moved