IRIS Intelligent Retrieval & Indexing System

A query in.
A timestamp out.

gradio space · api_name: run · sqlite index · runs locally

Retrieval over long recordings, scored down to the word. The whole system exists to answer one question well: at which second does this happen. Everything else is in service of that.

01 Stack

The only one of the four that can run entirely on your own machine, which is the point when the recording should not leave it.

Speechopenai/whisper-large-v3-turbo with word-level timestamps, not segment-level
Embeddingssentence-transformers/all-MiniLM-L6-v2
Framesopenai/clip-vit-large-patch14, sampled every 2 s, optional
IndexSQLite on disk. No vector service, no network dependency, no per-query cost.
Deviceauto. CUDA when present, CPU otherwise, same code path.
ServingGradio Space on free-tier hardware, or entirely local

02 How a query resolves

Indexing happens once per recording. Everything below that line runs per query, which is why the scoring is cheap on purpose.

Chunking

Audio is cut into overlapping windows. The overlap exists so a phrase that straddles a boundary still lands whole inside at least one chunk.

CHUNK_LENGTH 24 s · OVERLAP 4 s

Word-level timestamps

Segment-level timing would put you within half a minute of the answer. Word-level puts you on it, which is the difference between a useful jump and another round of scrubbing.

ASR_TIMESTAMPS word

Frame sampling

Frames are embedded alongside the audio, so a query can match something that was shown and never said out loud. Optional, because it roughly doubles indexing time.

FRAME_INTERVAL 2 s · VISUAL_SEARCH enabled

Hybrid scoring

Semantic similarity carries most of the weight, with a keyword component kept deliberately alive. Pure embeddings are weak on names, numbers and jargon, which is exactly what people search recordings for.

SEMANTIC 0.78 · KEYWORD 0.22 · KEYWORD_BOOST 0.15

Thresholds

Below these, a result is not returned at all. Returning the best of a bad set is how a search tool teaches people not to trust it.

MIN_SPEECH_SCORE 0.55 · MIN_VISUAL_SCORE 0.45

03 Decisions worth defending

Word timestamps over segmentCosts more to produce and more to store. Pays for itself the first time you land on the sentence instead of near it.
Hybrid, not pure vectorAn embedding model will happily rank a paragraph about deadlines above the one sentence containing the actual date. The keyword term is there to stop that.
SQLite, not a vector databaseOne file, no service to run, no bill per query, and it moves with the machine. At this scale the index fits and the queries are fast.
Returns nothing rather than somethingA floor under the score means an honest empty result. A tool that always answers is a tool nobody can calibrate against.