Google unveiled Gemini 3.5 Transcribe at Google I/O on May 19, 2026, a model capable of processing audio files within a 96,000-token context window. Beyond standard transcription, the model offers speaker identification, emotion detection, translation, and automatic summarization of audio content. The Gemini macOS application has integrated these features since mid-2026, while Google directs users to its Cloud Speech-to-Text API for real-time transcription. This solution targets professionals generating large volumes of spoken content: journalists, lawyers, academics, and medical practitioners. Competitive implications are significant for players like OpenAI and its Whisper tools, as well as dedicated transcription platforms.
Source: Read the original article

