Modulate Company Spotlight: AI Beyond the Transcript
Modulate is building artificial intelligence for the part of a conversation that disappears when speech becomes text. Words matter, but so do tone, hesitation, stress, timing, background sound, and whether the speaker is even human. The Somerville, Massachusetts company is turning those signals into an audio-intelligence layer for developers and enterprises.
That puts Modulate in a different lane from a conventional transcription vendor. Its platform is designed to help a system understand what is happening inside a voice interaction, not simply produce a written record after the fact. The company began with moderation in online games, where fast, noisy, emotionally charged speech created a difficult test, and is now carrying that production experience into fraud detection, customer operations, AI-agent oversight, and other enterprise workflows.
About Modulate
Modulate was founded in 2017 by MIT classmates Carter Huffman and Mike Pappas. Huffman is the current CEO, while Pappas serves as chairman. The company's team page traces the idea back to the pair's interest in how computers shape human interaction and Huffman's work on optimized machine-learning models at NASA's Jet Propulsion Laboratory.
Its first major product was ToxMod, a real-time voice-moderation system built for online games. Publicly named deployments include Call of Duty, Grand Theft Auto Online, Rainbow Six Siege, and Rec Room. Those environments gave Modulate a demanding operating ground: conversations move quickly, slang changes, microphones are messy, and the difference between banter and harm often depends on more than the literal words.
That history matters because it forced the company to work on live audio before voice agents and synthetic speech became enterprise priorities. It also gives Modulate credible production experience without proving that the same performance will automatically transfer to every industry or customer dataset.
Why Audio-Native Intelligence Matters Now
Most voice-analysis stacks begin by converting audio to text and sending the transcript into another model. That can work for summarization or keyword extraction. It becomes less reliable when the decision depends on pressure, sarcasm, frustration, vocal manipulation, overlapping speakers, or a synthetic voice.
Modulate's argument is that the original signal should remain central. Its Velma platform uses an Ensemble Listening Model, or ELM, that coordinates more than 100 specialized models. Different components examine speech, emotion, intent, speaker behavior, synthetic audio, language, background events, and other signals before an orchestration layer assembles them into a time-aligned interpretation.
The approach is deliberately modular. A team can inspect which signals contributed to a result and add or update specialized models as risks change. That does not eliminate model error, privacy questions, or the need for human review. It creates a different foundation for systems that need evidence from the sound itself.
From ToxMod to Velma
ToxMod remains the clearest example of Modulate's production origin. The system helps platforms surface potentially harmful voice interactions for review and action. The broader Velma platform extends the same audio-native architecture into developer APIs for transcription, conversation analysis, deepfake detection, emotion and accent detection, language and audio-event detection, AI-music detection, and PII or PHI redaction.
That expansion changes the company's addressable problem. A contact center may need to spot frustration before a customer churns. A financial institution may need evidence that a caller is using synthetic speech or social-engineering tactics. A company deploying voice agents may need to know when an automated system is confused, off-script, or creating risk.
These applications demand more than an impressive demo. They require low latency, predictable cost, explainability, privacy controls, integration with existing workflows, and strong performance on a customer's real audio. Modulate reports leading public benchmark results, including a 98.9% deepfake-detection result, but buyers still have to test the system against their languages, environments, attack patterns, and operating thresholds.
Funding Expands the Developer Bet
On September 28, 2026, Modulate announced $25M in new funding led by Future Ventures, with Hyperplane and Lakestar participating. The company says the financing brings total funding to $60M. It did not disclose a round-series label, valuation, ownership terms, or board changes.
Modulate plans to invest in AI and machine-learning research, product and engineering, developer relations, APIs, SDKs, models, deployment options, and partnerships. The practical goal is to make audio intelligence easier to add to products without asking every team to assemble separate systems for transcription, deepfake detection, emotion, moderation, and behavioral risk.
The company reports that it now analyzes more than 10M audio hours each month and has processed more than 600M hours in total. Those figures show operating scale, but they do not reveal revenue, retention, customer concentration, or audited impact. The next proof point is whether developers and enterprises adopt Velma as reusable infrastructure beyond the gaming base that established Modulate's reputation.
Hiring Shows the Operating Work Ahead
Modulate's careers page describes a hybrid team centered on collaboration in Somerville and lists a bias-reduction process that anonymizes application materials. The company also says it uses its own voice-masking technology and name anonymization during an early interview. Those practices offer a concrete example of the product informing internal operations, though they remain company-described rather than independently audited.
Hiring is a market signal because the next stage requires more than research talent. Modulate must package models into dependable products, support integrations, help developers evaluate outputs, and prove performance in customer-specific workflows. Product, engineering, developer-relations, and deployment work will decide whether an architecture becomes infrastructure.
What Modulate Signals for Voice AI
Voice is becoming both an interface and an attack surface. AI agents are making and receiving calls, synthetic speech is getting cheaper, and regulators are treating artificial voice as a material consumer-protection issue. The FTC has emphasized that voice-cloning harm has no single technical solution, while the FCC has applied artificial-voice rules to AI-generated robocalls.
Modulate is not the entire answer to that challenge. Its opportunity is to become the listening layer that gives other systems better evidence before they decide. If the company can preserve accuracy, speed, transparency, and cost advantages as it expands, audio-native intelligence could become a standard component in the voice stack rather than a specialized moderation tool.
The larger bet is simple to state and difficult to execute: a conversation contains more information than its transcript. Modulate has spent nearly a decade building around that premise. The new capital gives the team room to show how much of the voice economy is ready to listen the same way.
Frequently Asked Questions
What does Modulate do?
Modulate builds audio-native voice intelligence for developers and enterprises. Its platform analyzes raw audio for speech, emotion, intent, synthetic voice, behavioral risk, and other signals that a transcript alone may lose.
Who founded Modulate?
Carter Huffman and Mike Pappas founded Modulate in 2017 after meeting at MIT. Huffman is the current CEO, and Pappas serves as chairman.
What is Modulate's Ensemble Listening Model?
An Ensemble Listening Model coordinates more than 100 specialized models that examine different parts of an audio conversation, then combines their time-aligned signals into an explainable interpretation.
How are ToxMod and Velma different?
ToxMod is Modulate's real-time voice-moderation product, originally built for online games. Velma is the broader audio-intelligence platform and API for transcription, conversation analysis, deepfake detection, emotion, audio events, redaction, and related workflows.
How much funding has Modulate raised?
Modulate announced $25 million in new funding on September 28, 2026, led by Future Ventures with Hyperplane and Lakestar participating. The company says the round brings total funding to $60 million.
Where the Money Moved
The intelligence briefing of the innovation economy. Funding, M&A, debt and fund closes, read as market signal rather than deal announcements.
Subscribe to Where the Money Moved