How audio transcription with timestamps and event tagging works in Scribe
AIA native word-level transcription model outputs structured, timestamped arrays of word, spacing, and audio_event tokens directly from audio input, without a secondary forced-alignment pass. Audio events such as laughter or applause are tagged separately, which the source says helps with captioning, searchable archives, and highlight identification. The source notes Scribe's word-level transcription supports up to 5 independently transcribed channels.