Transcription and Text Extraction

Upload a video and Leed transcribes it. Upload an MP3 and Leed transcribes that too. Upload a PDF or a Word document and Leed pulls the words out of it. None of that is something you switch on, and none of it costs you a click: it starts the moment the upload finishes and writes its result onto the asset. What your plan controls is not whether the work happens — it is whether you can read the result.

What gets processed, and when

Asset typeTriggerWhat runsWhat it writes to the assetRuns on Free
Image—Nothing. Images get AI alt text instead, written on uploadaltTextYes
AudioThe moment the asset is createdTranscription (Whisper, on Cloudflare Workers AI)transcription, hasTranscriptionYes
VideoWhen the upload is finalized — then it waits for the encode to finish firstTranscription of the encoded audio tracktranscription, hasTranscriptionYes
DocumentThe moment the asset is created, and only for .pdf, .doc and .docxText extractiontext, hasExtractedTextYes

A document with any other extension uploads perfectly well and is served perfectly well — it is simply never queued for extraction, so its Extracted Content panel stays empty forever. Images are the other exception in the table and the one people ask about: nothing transcribes or extracts from a picture, and the machine-written description you get instead is covered in Uploading Images.

flowchart TD
  U["You upload an asset"] --> K{"What kind?"}
  K -->|"PDF / DOC / DOCX"| DX["Document extraction"]
  K -->|"Other document type"| NONE["Nothing runs"]
  K -->|"Audio"| AX["Audio transcription"]
  K -->|"Video"| W["Wait for the encode to finish"]
  W --> VX["Video transcription"]
  DX --> S["Result stored on the asset<br/>on EVERY plan"]
  AX --> S
  VX --> S
  S --> G{"Your plan"}
  G -->|"Starter and up"| SHOW["The text is returned<br/>panel shows it, Insert is enabled"]
  G -->|"Free"| HIDE["The field is removed from the response<br/>panel shows an upgrade banner, Insert is disabled"]

Where the result appears

Both results live on the asset itself, in a collapsed section near the bottom of the asset page. Open Design → Assets, click the asset, and expand the section.

A video asset in the CMS with its Transcript section expanded, showing the transcribed text

The section is called Transcript and sits between Player Settings and File Details. On a workspace with the feature it shows the transcript text, or No transcript if the video has been processed and nothing came back. Note that the video becomes playable before the transcript exists — the same workflow does both, and it marks the video ready first.

Everything else on the asset page — the poster frame, the playback checkboxes, the filename, the PDF preview — is covered in Video, Audio and Document Assets.

Putting the text on a page

You do not have to copy and paste. When you place a video, an audio file or a document in the editor, the block’s settings dialog carries a button that drops the text into the page as real editor blocks below the media — paragraphs you can then edit, cut and reformat like anything else you typed.

  • On a video or audio block, the button is Insert transcript.
  • On the document picker, it is Insert extracted text.

The button is disabled in two different situations, and its tooltip tells you which one you are in. These are the three states, verbatim, from a video block:

Auto-transcription & content extraction is available on the Starter plan
Insert the transcript as editor blocks below this video
No transcript available yet — transcripts are generated after the asset is processed

An audio block says “below this audio” instead of “below this video”; the document picker says “Insert the extracted text as editor blocks below this paragraph” and “No extracted text available yet — text is extracted after the document is processed”. The first line is the plan gate; the third is simply “not finished yet, or never queued”.

In the block’s settings dialog the control sits below the playback checkboxes and reads Insert transcript. It is enabled only when the asset actually carries transcript text — otherwise it is greyed out and the dialog says no transcript is available. Hovering it shows the tooltip “Insert the transcript as editor blocks below this video”.

Where the media dialogs live and how you get to them is on Inserting Media and Embeds.

How long it takes, and how you know

Rough envelopes, so you know when to stop waiting:

Asset typeWhat has to happenPractical wait
DocumentOne conversion pass, with a five-minute budget and up to two retriesSeconds to a couple of minutes
AudioOne transcription pass, with a ten-minute budget and up to two retriesUnder a minute for a short clip
VideoThe encode first — Leed checks with the video platform every 30 seconds, up to 120 times — then transcriptionMinutes; the ceiling on the wait for the encode alone is about an hour

Video transcription works on the encoded stream rather than your original file: Leed groups the stream into roughly 30-second batches with a 3-second overlap between neighbors, so a word spoken across a boundary is not lost, and transcribes the batches in parallel. Long videos are therefore not much slower than short ones once the encode is done.

Nothing appeared after an hour

Work through these in order.

  1. Reload the asset page. The panel does not refresh on its own, and this is by far the most common answer.
  2. Check the file type. Only .pdf, .doc and .docx are ever queued for extraction. A .txt, .pptx, .csv or .zip file will never produce extracted content, and that is by design rather than a failure.
  3. Check the document has a text layer. A PDF that is a photograph of a page contains no text to extract, and comes back as No content extracted.
  4. For a video, check whether it is still showing “Processing…”. If it is, the encode has not finished or has failed — and because the transcription workflow is what marks a video ready, a video whose encode genuinely failed can sit on Processing… indefinitely. That case, and what to do about it, is in Known Limitations.
  5. If none of those apply, re-uploading the file starts a fresh run. Send support the asset’s name and the time you uploaded it.

What the text is, and what it is not

Set your expectations from the machine, not from a court reporter.

A transcript is one continuous block of machine-written text. It has no speaker labels, no timestamps and no guarantee about punctuation or the spelling of names. It is not a caption file — there is no WebVTT track, nothing is attached to the player, and turning on captions is not something this feature does. There is also no editing surface: you cannot correct a transcript on the asset. If you want a corrected version on a page, insert it and edit the resulting blocks.

Extracted document text is stripped down harder than most people expect. Everything outside letters, digits, whitespace, hyphens and the four marks , . ! ? is removed before it is stored. In practice that means tables lose their structure, currency and percentage symbols disappear, quotation marks and brackets go, and accented characters are dropped rather than transliterated — café becomes caf. The extraction is not broken when you see this; it is doing exactly what it is built to do, which is produce plain searchable prose rather than a faithful reproduction of your document.

What the Starter plan unlocks

The processing is free. The retrieval is what contentExtraction gates, and it gates it by removing fields from a response rather than by refusing the request.

SurfaceStarter and upFree
Asset Transcript panel (video, audio)The transcript textAn upgrade banner in the same place
Asset Extracted Content panel (document)The extracted textAn upgrade banner in the same place
Editor Insert transcript buttonEnabled once a transcript existsDisabled, with the upgrade line as its tooltip
Editor Insert extracted text buttonEnabled once text has been extractedDisabled, with the upgrade line as its tooltip
Fetching a single asset over the APIThe asset includes transcription or textThose fields are absent, and the response carries contentExtractionGated: true
Listing assets, over the API or over MCPNo transcript in the listIdentical — the list has never carried the text on any plan
The “a transcript exists” flaghasTranscription / hasExtractedText presentIdentical — the flag is never stripped

That last row is the one that makes the whole design work. Because the existence flag survives the gate, a Free workspace can be told this asset has a transcript, upgrade to read it rather than the much worse and less true no transcript found.

The same Transcript section on a Free workspace, showing the upgrade banner in place of the text

The banner reads “Auto-transcription & content extraction is available on the Starter plan”, and offers an Upgrade plan button to anyone with billing access — or, to everyone else, the line “Ask your account admin to upgrade your plan.”

Why capture and then gate

Leed could have skipped the processing on Free and saved itself the work. It does not, and the reason is worth one sentence: a workspace that upgrades on Tuesday gets every transcript and every extracted document it has ever uploaded on Tuesday afternoon, because the text was captured all along. The alternative — processing only from the upgrade onwards — would hand you a feature that works on new uploads and silently fails on your archive.

What Leed never does with it

The text sits on the asset record and goes nowhere on its own. It is not published to your site, not written into your page bodies, not added to your site’s search index and not sent anywhere else. The only thing that ever moves it is you, pressing Insert transcript or Insert extracted text — and from that moment it is ordinary page content, indistinguishable from words you typed.

Transcription is the one AI surface in Leed that runs without you asking for it; the others, and which one you actually want, are indexed in AI in Leed.

ESC