NINA Neural Information & Narrative Abstraction

Recording in.
Time-anchored structure out.

gradio space · api_name: run / chat / route / fetch

Long audio and video turned into structure that still knows what time it is: transcript, speaker turns, semantic segments, an index, and a synthesis that can point at the segments it used. This page is the build, not the pitch.

01 Stack

Everything here is an open checkpoint. Nothing in the chain depends on a single vendor's endpoint staying up or staying priced the way it is today.

Speechwhisper base for the first pass, whisper large-v3-turbo for retranscription
Speakerspyannote/speaker-diarization-3.1, best effort: skipped when the token lacks access
Framesclip-ViT-B-32, up to 96 frames sampled at 0.1 to 0.5 fps, 224 px
SynthesisQwen3-VL-30B-A3B-Instruct, text-only fallback Qwen2.5-72B-Instruct
ServingGradio on Hugging Face Spaces, driven server-side from a Cloudflare Worker
AuthHMAC-SHA256 token, 120 s TTL, minted per use. The Space is public, the model is not.

02 Pipeline

Seven stages. The interesting ones are the two that exist purely to stop the later stages from confidently building on something wrong.

Ingest

A direct upload, or a pasted link fetched server-side. Oversized video is re-encoded before upload, because the model needs audio and sparse frames, not resolution.

audio 16 kHz mono · video downscaled above the trigger size

Transcription, first pass

The fast model runs over everything. It is good enough for most speech and cheap enough to run on the whole file.

WHISPER_FAST base · CPU_THREADS 3 · CONDITION_ON_PREV False

Transcription, second pass

Only the segments whose confidence fell below the gate are transcribed again with the precise model. Running the expensive model over everything costs far more than it returns; running it where the fast one wobbled costs a fraction and fixes the parts that were actually wrong.

PRECISE_LOGPROB_GATE −0.40 · WHISPER_PRECISE large-v3-turbo · RETRANSCRIBE_MAX 12

Reliability gate

The transcript is scored before anything is built on it. A recording that is mostly noise gets reported as unreliable rather than summarised with confidence. A tidy summary of a bad transcript is worse than no summary, because nobody checks it.

LOGPROB_FLOOR −0.85 · WEAK_FRACTION 0.5 · NO_SPEECH_CEILING 0.55

Segmentation

Cuts fall where the topic shifts, not on a fixed clock. Fixed windows split sentences in half and glue unrelated subjects together, and both hurt retrieval more than even spacing helps it.

SPLIT_THRESHOLD 0.45 · MIN 8 s · MAX 90 s · WINDOW 5

Retrieval

A bi-encoder pulls a shortlist, then a cross-encoder reranks it. The first is fast and approximate, the second slow and accurate. Forty then ten is where that trade stops being worth it.

BI_TOP_K 40 → CROSS_TOP_K 10 → RESULTS 6 · floors 0.10 global / 0.28 hybrid

Synthesis

The vision model reads the shortlisted segments with their timestamps still attached and writes the answer, which is why the answer can point back at a minute rather than a paragraph.

MAX_TOKENS 1800 · TEMPERATURE 0.05 · TIMEOUT 120 s · VISUAL_WEIGHT 0.30

03 Decisions worth defending

Four choices that cost something, and what they bought.

Two-pass ASRThe precise model over the whole file would be the obvious build. It is also mostly wasted: the fast model is already right where the audio is clean. Gating on confidence spends the budget only where it changes the output.
Gate before synthesisWithout it, a noisy recording still produces a fluent summary, and fluency reads as correctness. Failing loudly is the cheaper error.
Cut on meaningFixed-length chunks are simpler and retrieve worse. A segment that starts mid-sentence matches nothing well.
Turn runs server-sideThe Worker drives the Space through ctx.waitUntil, so closing the tab does not kill the answer. It lands in the database either way.