Full transcripts: availability and ingestion
No full transcript is available in this initial release. Availability is explicit in both talk pages and the machine-readable talk records.
Publication contract
Supply authorized timestamped speech text or nonempty original captions. Record speaker, title, event, language, recording URL, official event URL, transcription method, permission basis, retrieval date, correction history, and review status. Do not name OpenAI or another transcription provider unless it actually produced the text.
Publish the full transcript as transcript.md and transcript.txt, and actual timed captions as captions.vtt. An optional captions.srt must be converted from the same cues. Preserve the spoken language; label any translation separately.
Use timestamp headings in the Markdown transcript, linking to the original video at the corresponding number of seconds. Add machine-generated status and a warning to verify important quotations against the original recording. Correct technical transcription mistakes conservatively; do not paraphrase speech into a polished article.
The current renderer intentionally rejects YAML front matter. Before ingesting real transcripts, add explicit front-matter parsing and tests for their metadata, ordered time ranges, nonempty cues, and consistency across Markdown, plain text, VTT, and SRT. Then change transcript_status and add transcript, caption, and text URLs to the dataset and talk pages.
Until that validated ingestion path exists, do not replace an unavailable notice with unvalidated captions or fabricate timestamped text.