Check for speech text
The pipeline first asks for an existing caption or transcript track because it is faster and usually cleaner.
No caption track is not the end of the road
When a public video has no captions, the transcript provider can fall back to AI audio transcription. The summary is then built from that generated transcript—not guessed from the title.
This page deliberately does not substitute a captioned marketing demo. Paste a public video whose YouTube player has no “Show transcript” option; after AI audio transcription completes, the full Transcript tab and the resulting summary appear together.
The pipeline first asks for an existing caption or transcript track because it is faster and usually cleaner.
When the provider finds no caption track, its speech-recognition fallback can create a timed transcript from the audio.
The generated transcript feeds the same key-point, prose and chapter workflow and remains available for inspection.
A fluent answer is not proof that a tool heard the video.
Confirm the YouTube player itself has no transcript option before submitting the link.
A full timed transcript demonstrates that speech was retrieved or transcribed before the summary was written.
If neither source can produce text, the correct outcome is an error—not a plausible recap based on metadata.
Turning subtitle display off is different from a video having no caption track. If a track exists, the service can use it whether or not the viewer has captions visible.
No. The video must be public and its audio accessible and intelligible. Private, members-only, blocked, silent or severely distorted videos may not produce a transcript.
Open the transcript and timestamps. If no speech text can be produced, the API returns a no-transcript error instead of asking the model to infer a summary from the title.