Google AI has released Gemini 3.5 Transcribe, a speech-to-text model achieving 2.6% average Word Error Rate across 85+ languages. The release includes two variants: gemini-3.5-transcribe-live for real-time WebSocket streaming (no diarization, 10-minute session cap) and gemini-3.5-transcribe for pre-recorded files with diarization, word offsets, and custom vocabulary support up to
reddit/r/machinelearningnews