Skip to content

Transcribe Audio

Turn speech into text entirely in your browser with Whisper (tiny/base, on transformers.js). Private, offline, no cost — audio never leaves the device.

Runs the audio through an on-device speech-to-text model and returns the transcript plus timestamped segments. The audio is decoded and resampled in the browser before it reaches the model.

Transcribe a voice note, a recorded call, or microphone input captured by Record Audio — then summarize, search, or route the text downstream.

SettingNotes
ModelAny installed speech-to-text model (e.g. Whisper tiny/base).
AudioAudio URL, data URL, or $binary — pair with Record Audio or a file node. The Record tab lets you record a clip from your microphone right in the settings.

Next to From input, URL, Upload and Base64, the Audio field has a Record tab: record a clip (up to 2 minutes) from your microphone — you can pick which mic — and it’s stored with the node like an upload. Use it to try a model with Run preview, or as fixed audio for the node. To capture new audio each time the workflow runs, use the Record Audio node instead.

Returns { transcript, segments, chunks }:

  • transcript — the full text.
  • segments — timed pieces of the transcript as { start, end, text }, times in seconds. The last segment’s end can be null when the audio cuts off mid-phrase.
  • chunks — the same timing in the model’s raw form ({ timestamp: [start, end], text }), kept for older workflows.
  • Empty transcript — check the audio actually contains speech and is a format the browser can decode (e.g. webm, wav, mp3).
  • Slow first run — Whisper loads on first use, then stays warm for the session. Prefer tiny on low-memory devices.