Gemini 3.5 Transcribe brings emotion detection and speaker ID to speech-to-text

Share

Google unveiled Gemini 3.5 Transcribe at Google I/O on May 19, 2026, a model capable of processing audio files within a 96,000-token context window. Beyond standard transcription, the model offers speaker identification, emotion detection, translation, and automatic summarization of audio content. The Gemini macOS application has integrated these features since mid-2026, while Google directs users to its Cloud Speech-to-Text API for real-time transcription. This solution targets professionals generating large volumes of spoken content: journalists, lawyers, academics, and medical practitioners. Competitive implications are significant for players like OpenAI and its Whisper tools, as well as dedicated transcription platforms.

Source: Read the original article

Telemac
Telemachttp://cryptoinfo.ch
Passionné de nouvelles technologies, j’explore l’univers de la blockchain et des cryptomonnaies pour partager l’actualité et les innovations du secteur.

Lire la Suite

Articles