NINA Neural Information & Narrative Abstraction
Long audio and video turned into structure that still knows what time it is: transcript, speaker turns, semantic segments, an index, and a synthesis that can point at the segments it used. This page is the build, not the pitch.
Everything here is an open checkpoint. Nothing in the chain depends on a single vendor's endpoint staying up or staying priced the way it is today.
whisper base for the first pass, whisper large-v3-turbo for retranscriptionpyannote/speaker-diarization-3.1, best effort: skipped when the token lacks accessclip-ViT-B-32, up to 96 frames sampled at 0.1 to 0.5 fps, 224 pxQwen3-VL-30B-A3B-Instruct, text-only fallback Qwen2.5-72B-InstructSeven stages. The interesting ones are the two that exist purely to stop the later stages from confidently building on something wrong.
A direct upload, or a pasted link fetched server-side. Oversized video is re-encoded before upload, because the model needs audio and sparse frames, not resolution.
The fast model runs over everything. It is good enough for most speech and cheap enough to run on the whole file.
Only the segments whose confidence fell below the gate are transcribed again with the precise model. Running the expensive model over everything costs far more than it returns; running it where the fast one wobbled costs a fraction and fixes the parts that were actually wrong.
The transcript is scored before anything is built on it. A recording that is mostly noise gets reported as unreliable rather than summarised with confidence. A tidy summary of a bad transcript is worse than no summary, because nobody checks it.
Cuts fall where the topic shifts, not on a fixed clock. Fixed windows split sentences in half and glue unrelated subjects together, and both hurt retrieval more than even spacing helps it.
A bi-encoder pulls a shortlist, then a cross-encoder reranks it. The first is fast and approximate, the second slow and accurate. Forty then ten is where that trade stops being worth it.
The vision model reads the shortlisted segments with their timestamps still attached and writes the answer, which is why the answer can point back at a minute rather than a paragraph.
Four choices that cost something, and what they bought.
ctx.waitUntil, so closing the tab does not kill the answer. It lands in the database either way.