---
title: "Interviews, lectures and videos you can cite to the second"
description: "Word-by-word transcription with who is speaking, citable stretches of 30 to 60 seconds, searchable video frames, and a player with a synchronised transcript where you cite by selecting the text."
url: https://scholaris.joseluissaorin.com/en/knowledge/media-player
markdown: https://scholaris.joseluissaorin.com/en/knowledge/media-player.md
lang: en
alternate_es: https://scholaris.joseluissaorin.com/saber/reproductor.md
updated: 2026-10-06
author: José Luis Saorín Ferrer (https://joseluissaorin.com)
---

# Interviews, lectures and videos you can cite to the second

> Word-by-word transcription with who is speaking, citable stretches of 30 to 60 seconds, searchable video frames, and a player with a synchronised transcript where you cite by selecting the text.

## The transcript

Audio is transcribed in ten-minute stretches with a two-second overlap. When there is a Gemini key, Gemini Transcribe does it, telling speakers apart; if it fails, Whisper (large-v3-turbo, on Workers AI). Then a model identifies the cast (who speaks and in what role: interviewer, guest, host) and assigns every sentence to a named person, which fixes labels that do not match from one stretch to the next. In two interviews from the benchmark, attribution was right in 20 of 21 turns and in 23 of 23.

The text is cut into **citable stretches** of 30 to 60 seconds (45 on average) that end at a sentence boundary. Each stretch is a unit with its time anchor and speaker, and per-word timings are stored too.

## Frames

In videos a frame is taken at every scene change and, if nothing changes, one every 20 seconds (never more than one every 10). Frames are described and vectorised, so you can search for "the blackboard with the diagram" and land on the exact second.

## YouTube, Vimeo and podcasts

A YouTube video is not downloaded: Gemini watches it from its address, ten minutes at a time, and the title and channel come from its public record. From Vimeo the smallest downloadable file is taken; from a podcast, the audio its RSS announces.

## The player

- A timeline with each speaker's turns, the chapters and a frame preview on hover.
- A synchronised transcript that highlights the word being spoken; click a word to jump there.
- To cite, select the text: the citation comes with its exact interval ("12:04-12:40").
- A small player keeps playing while you move around the rest of the app.

## What we measured

A 54-minute video interview was ready in 45 seconds ($0.41), and a two-hour one in 70 seconds ($0.94); transcription costs about $0.005 a minute. In the first Scholaris, the same 54-minute interview took three hours and 39 minutes. All figures and their conditions are in [Performance](https://scholaris.joseluissaorin.com/en/knowledge/performance.md).

## What is missing

Per-word timings are estimated within each segment, not measured one by one. Some videos with uncommon codecs are not yet converted to a format every browser can play.
