IRIS Intelligent Retrieval & Indexing System
Retrieval over long recordings, scored down to the word. The whole system exists to answer one question well: at which second does this happen. Everything else is in service of that.
The only one of the four that can run entirely on your own machine, which is the point when the recording should not leave it.
openai/whisper-large-v3-turbo with word-level timestamps, not segment-levelsentence-transformers/all-MiniLM-L6-v2openai/clip-vit-large-patch14, sampled every 2 s, optionalauto. CUDA when present, CPU otherwise, same code path.Indexing happens once per recording. Everything below that line runs per query, which is why the scoring is cheap on purpose.
Audio is cut into overlapping windows. The overlap exists so a phrase that straddles a boundary still lands whole inside at least one chunk.
Segment-level timing would put you within half a minute of the answer. Word-level puts you on it, which is the difference between a useful jump and another round of scrubbing.
Frames are embedded alongside the audio, so a query can match something that was shown and never said out loud. Optional, because it roughly doubles indexing time.
Semantic similarity carries most of the weight, with a keyword component kept deliberately alive. Pure embeddings are weak on names, numbers and jargon, which is exactly what people search recordings for.
Below these, a result is not returned at all. Returning the best of a bad set is how a search tool teaches people not to trust it.