Monday, August 3, 2026

Apps & Consumer

DeepL launches voice-to-voice translation suite

DeepL has launched a voice-to-voice translation suite and developer API, with plans to eventually develop an end-to-end model that skips the text-translation step entirely.

DeepL launches voice-to-voice translation suite
Photo: DeepL

DeepL, a translation company known for its text tools, released a voice-to-voice translation suite today. The new product covers use cases such as meetings, mobile and web conversations, and group conversations. For group settings like training sessions or workshops, participants can join the conversation through a QR code. As part of this launch, DeepL is releasing add-ons for platforms like Zoom and Microsoft Teams, allowing users to either hear real-time translation or follow translated text on-screen. These integrations are currently in early access, with the company inviting organizations to join a waitlist.

The company is also releasing an Application Programming Interface (API) that lets outside developers and businesses build on top of the technology for use cases such as call centers. The voice-to-voice technology is designed to adapt to custom vocabulary, including industry-specific terms and company or personal names. Currently, the system operates by converting speech to text, translating it, and then converting it back to speech. However, DeepL wants to develop an end-to-end voice translation model that skips the text step entirely.

The expansion into voice represents a natural evolution for the company. “After spending so many years in text translation, voice was a natural step for us,” DeepL CEO Jarek Kutylowski said. He noted that the primary technical challenge in creating a real-time translation product centers on balancing low latency—the delay between someone speaking and the translated audio playing back—with accurate results. Kutylowski also highlighted that a translation layer helps companies provide customer support in languages where qualified staff are scarce and expensive to hire.

DeepL enters a competitive landscape populated by several startups targeting different niches within speech synthesis and real-time translation:

  • Sanas: This startup uses AI to modify a speaker’s accent in real time, primarily for call center agents. Sanas raised $65 million last year from Quadrille Capital, with participation from Teleperformance.
  • Camb.AI: Based in Dubai, this company focuses on speech synthesis and translation for media and entertainment companies to help them dub and localize video content.
  • Palabra: Backed by Alexis Ohanian’s firm Seven Seven Six, this startup is building a real-time speech translation engine designed to preserve both the meaning and the speaker’s original voice.

Why it matters

DeepL is expanding its translation capabilities from text to real-time voice to capture enterprise use cases like meetings and customer service, positioning itself against a wave of specialized startups.