KM krzysztof mirecki

Retrieval.
Down to the second.

iris · whisper large-v3-turbo · minilm · clip vit-l/14 · sqlite

One question, answered properly: at which second does this happen. Everything in IRIS is in service of that, including the parts that refuse to answer.

01 What I work on

IRISIndexing and retrieval over long recordings, word-level, with an optional visual channel
ScoringHybrid semantic and keyword, because pure embeddings lose exactly the names and numbers people search for
StorageSQLite on disk. One file, no service, no per-query bill, and it moves with the machine.
DeploymentRuns on free-tier hardware or entirely locally, same code path

02 Problems I spent the most time on

Landing on it, not near it

Segment-level timestamps put you within half a minute of the answer, which still means scrubbing. Word-level timing costs more to produce and store, and removes the last step.

ASR_TIMESTAMPS word

Chunks that do not cut sentences

Fixed windows split a phrase across a boundary and then neither half matches well. Overlap means every phrase lands whole in at least one chunk.

CHUNK 24 s · OVERLAP 4 s

Keeping keywords alive

An embedding model ranks a paragraph about deadlines above the one sentence with the actual date. A keyword term stops that, weighted so it corrects rather than dominates.

SEMANTIC 0.78 · KEYWORD 0.22 · BOOST 0.15

Saying nothing

A search tool that always returns its best guess cannot be calibrated against. Below the floor, IRIS returns an empty result and says so.

MIN_SPEECH 0.55 · MIN_VISUAL 0.45

03 Stack

Speechopenai/whisper-large-v3-turbo, word timestamps
Embeddingssentence-transformers/all-MiniLM-L6-v2
Visionopenai/clip-vit-large-patch14, frames every 2 s
RuntimeFastAPI, SQLite, device auto so CUDA and CPU share one path