AI · 3 Oktober 2026
Modulate Raises $25M to Scale Audio-Native AI Models
Audio-native AI company Modulate has raised $25 million led by Future Ventures, bringing its total funding to $60 million, to scale its voice-first AI models and developer ecosystem.
What happened
Audio-native AI company Modulate has closed a $25 million funding round led by Future Ventures, taking its total capital raised to $60 million. The company says the new funding will be used to scale its frontier audio-native AI models and grow the developer ecosystem built around them.
The raise signals continued investor appetite for AI systems built specifically around audio and voice data, rather than text-first models adapted after the fact to handle speech.
Why it matters
Most generative AI investment to date has concentrated on text and, increasingly, multimodal vision-language systems. A dedicated push into audio-native modelling points to a maturing recognition that voice carries information — tone, intent, emotion, urgency — that text transcription alone discards. Funding rounds of this size for a specialist audio AI player suggest the category is being treated as its own frontier, not a feature bolted onto existing language models.
For organisations building voice-based products — contact centres, gaming platforms, voice assistants, trust-and-safety tooling — the emergence of better-funded, more capable audio-native models expands what's technically possible: finer-grained understanding of spoken interactions, and potentially faster, cheaper ways to build voice applications without routing everything through a text intermediary first. A broader developer ecosystem around these models also lowers the barrier for smaller players to experiment with voice AI rather than relying solely on the handful of dominant foundation model providers.
By the numbers
- $25 million raised in the new funding round, led by Future Ventures
- $60 million total funding raised by Modulate to date
The Renascence take
It's tempting to file this as just another AI funding story, but the audio-native framing is the detail worth sitting with. Voice is the oldest and still the most emotionally dense channel in customer and user experience, and most AI tooling has treated it as an afterthought — a transcript to be text-mined rather than a signal in its own right.
The real opportunity in audio-native AI isn't faster transcription — it's models that treat hesitation, tone and pacing as data, not noise. Organisations still routing voice interactions through a text pipeline before analysing them are discarding the richest signal in the conversation. As this category attracts serious capital, the operators who benefit first won't be the ones chasing the flashiest voice bot, but those who use audio-native understanding to catch friction and emotion that transcripts have always missed.
Sumber-sumber
Taklimat ini ditulis oleh Meja Berita kami, mensintesis laporan daripada saluran di bawah. Ikuti pautan untuk liputan asal.
FAQ
Questions we get on this topic
Lagi dalam AI
Kekal di hadapan CX
Dapatkan isyarat, bukan gangguan.
Kisah-kisah yang membentuk pengalaman pelanggan — serta Jurnal dan Experience Loom — di peti masuk anda.
