Upload a video and Leed transcribes it. Upload an MP3 and Leed transcribes that too. Upload a PDF or a Word document and Leed pulls the words out of it. None of that is something you switch on, and none of it costs you a click: it starts the moment the upload finishes and writes its result onto the asset. What your plan controls is not whether the work happens — it is whether you can read the result.
What gets processed, and when
| Asset type | Trigger | What runs | What it writes to the asset | Runs on Free |
|---|---|---|---|---|
| Image | — | Nothing. Images get AI alt text instead, written on upload | altText | Yes |
| Audio | The moment the asset is created | Transcription (Whisper, on Cloudflare Workers AI) | transcription, hasTranscription | Yes |
| Video | When the upload is finalized — then it waits for the encode to finish first | Transcription of the encoded audio track | transcription, hasTranscription | Yes |
| Document | The moment the asset is created, and only for .pdf, .doc and .docx | Text extraction | text, hasExtractedText | Yes |
A document with any other extension uploads perfectly well and is served perfectly well — it is simply never queued for extraction, so its Extracted Content panel stays empty forever. Images are the other exception in the table and the one people ask about: nothing transcribes or extracts from a picture, and the machine-written description you get instead is covered in Uploading Images.
flowchart TD
U["You upload an asset"] --> K{"What kind?"}
K -->|"PDF / DOC / DOCX"| DX["Document extraction"]
K -->|"Other document type"| NONE["Nothing runs"]
K -->|"Audio"| AX["Audio transcription"]
K -->|"Video"| W["Wait for the encode to finish"]
W --> VX["Video transcription"]
DX --> S["Result stored on the asset<br/>on EVERY plan"]
AX --> S
VX --> S
S --> G{"Your plan"}
G -->|"Starter and up"| SHOW["The text is returned<br/>panel shows it, Insert is enabled"]
G -->|"Free"| HIDE["The field is removed from the response<br/>panel shows an upgrade banner, Insert is disabled"]
Where the result appears
Both results live on the asset itself, in a collapsed section near the bottom of the asset page. Open Design → Assets, click the asset, and expand the section.
- Video
- Audio
- Document
The section is called Transcript and sits between Player Settings and File Details. On a workspace with the feature it shows the transcript text, or No transcript if the video has been processed and nothing came back. Note that the video becomes playable before the transcript exists — the same workflow does both, and it marks the video ready first.
The section is called Transcript and sits between Player Settings and File Details, exactly as it does for video. Audio has no encode step to wait for, so a short MP3’s transcript is usually there by the time you have finished naming the file.
The section is called Extracted Content and sits directly above File Details. It shows the extracted text, or No content extracted for a document that was processed and yielded nothing — which is what you get from a scanned PDF with no text layer.
Everything else on the asset page — the poster frame, the playback checkboxes, the filename, the PDF preview — is covered in Video, Audio and Document Assets.
Putting the text on a page
You do not have to copy and paste. When you place a video, an audio file or a document in the editor, the block’s settings dialog carries a button that drops the text into the page as real editor blocks below the media — paragraphs you can then edit, cut and reformat like anything else you typed.
- On a video or audio block, the button is Insert transcript.
- On the document picker, it is Insert extracted text.
The button is disabled in two different situations, and its tooltip tells you which one you are in. These are the three states, verbatim, from a video block:
Auto-transcription & content extraction is available on the Starter plan
Insert the transcript as editor blocks below this video
No transcript available yet — transcripts are generated after the asset is processedAn audio block says “below this audio” instead of “below this video”; the document picker says “Insert the extracted text as editor blocks below this paragraph” and “No extracted text available yet — text is extracted after the document is processed”. The first line is the plan gate; the third is simply “not finished yet, or never queued”.
In the block’s settings dialog the control sits below the playback checkboxes and reads Insert transcript. It is enabled only when the asset actually carries transcript text — otherwise it is greyed out and the dialog says no transcript is available. Hovering it shows the tooltip “Insert the transcript as editor blocks below this video”.
Where the media dialogs live and how you get to them is on Inserting Media and Embeds.
How long it takes, and how you know
Rough envelopes, so you know when to stop waiting:
| Asset type | What has to happen | Practical wait |
|---|---|---|
| Document | One conversion pass, with a five-minute budget and up to two retries | Seconds to a couple of minutes |
| Audio | One transcription pass, with a ten-minute budget and up to two retries | Under a minute for a short clip |
| Video | The encode first — Leed checks with the video platform every 30 seconds, up to 120 times — then transcription | Minutes; the ceiling on the wait for the encode alone is about an hour |
Video transcription works on the encoded stream rather than your original file: Leed groups the stream into roughly 30-second batches with a 3-second overlap between neighbors, so a word spoken across a boundary is not lost, and transcribes the batches in parallel. Long videos are therefore not much slower than short ones once the encode is done.
Nothing appeared after an hour
Work through these in order.
- Reload the asset page. The panel does not refresh on its own, and this is by far the most common answer.
- Check the file type. Only
.pdf,.docand.docxare ever queued for extraction. A.txt,.pptx,.csvor.zipfile will never produce extracted content, and that is by design rather than a failure. - Check the document has a text layer. A PDF that is a photograph of a page contains no text to extract, and comes back as
No content extracted. - For a video, check whether it is still showing “Processing…”. If it is, the encode has not finished or has failed — and because the transcription workflow is what marks a video ready, a video whose encode genuinely failed can sit on Processing… indefinitely. That case, and what to do about it, is in Known Limitations.
- If none of those apply, re-uploading the file starts a fresh run. Send support the asset’s name and the time you uploaded it.
What the text is, and what it is not
Set your expectations from the machine, not from a court reporter.
A transcript is one continuous block of machine-written text. It has no speaker labels, no timestamps and no guarantee about punctuation or the spelling of names. It is not a caption file — there is no WebVTT track, nothing is attached to the player, and turning on captions is not something this feature does. There is also no editing surface: you cannot correct a transcript on the asset. If you want a corrected version on a page, insert it and edit the resulting blocks.
Extracted document text is stripped down harder than most people expect. Everything outside letters, digits, whitespace, hyphens and the four marks , . ! ? is removed before it is stored. In practice that means tables lose their structure, currency and percentage symbols disappear, quotation marks and brackets go, and accented characters are dropped rather than transliterated — café becomes caf. The extraction is not broken when you see this; it is doing exactly what it is built to do, which is produce plain searchable prose rather than a faithful reproduction of your document.
What the Starter plan unlocks
The processing is free. The retrieval is what contentExtraction gates, and it gates it by removing fields from a response rather than by refusing the request.
| Surface | Starter and up | Free |
|---|---|---|
| Asset Transcript panel (video, audio) | The transcript text | An upgrade banner in the same place |
| Asset Extracted Content panel (document) | The extracted text | An upgrade banner in the same place |
| Editor Insert transcript button | Enabled once a transcript exists | Disabled, with the upgrade line as its tooltip |
| Editor Insert extracted text button | Enabled once text has been extracted | Disabled, with the upgrade line as its tooltip |
| Fetching a single asset over the API | The asset includes transcription or text | Those fields are absent, and the response carries contentExtractionGated: true |
| Listing assets, over the API or over MCP | No transcript in the list | Identical — the list has never carried the text on any plan |
| The “a transcript exists” flag | hasTranscription / hasExtractedText present | Identical — the flag is never stripped |
That last row is the one that makes the whole design work. Because the existence flag survives the gate, a Free workspace can be told this asset has a transcript, upgrade to read it rather than the much worse and less true no transcript found.
The banner reads “Auto-transcription & content extraction is available on the Starter plan”, and offers an Upgrade plan button to anyone with billing access — or, to everyone else, the line “Ask your account admin to upgrade your plan.”
Why capture and then gate
Leed could have skipped the processing on Free and saved itself the work. It does not, and the reason is worth one sentence: a workspace that upgrades on Tuesday gets every transcript and every extracted document it has ever uploaded on Tuesday afternoon, because the text was captured all along. The alternative — processing only from the upgrade onwards — would hand you a feature that works on new uploads and silently fails on your archive.
What Leed never does with it
The text sits on the asset record and goes nowhere on its own. It is not published to your site, not written into your page bodies, not added to your site’s search index and not sent anywhere else. The only thing that ever moves it is you, pressing Insert transcript or Insert extracted text — and from that moment it is ordinary page content, indistinguishable from words you typed.
Transcription is the one AI surface in Leed that runs without you asking for it; the others, and which one you actually want, are indexed in AI in Leed.