Whisper + forced phoneme alignment + speaker diarization. Word-level timestamps and who-said-what tagging for any audio. Free and open.