LUNA Latent Understanding & Narrative Assembly
A hundred mixed files do not fit in any context window, so LUNA does not try. It reads each one separately with a cheap model, keeps only what the brief needs, and hands the survivors to an expensive one.
Two language models with different jobs and very different price tags, plus a vision model for anything that arrives as a picture.
Qwen/Qwen2.5-7B-Instruct, temperature 0.0, 900 tokens per file, 8 workers in parallelQwen/Qwen2.5-72B-Instruct, temperature 0.15, up to 4000 tokensQwen/Qwen3-VL-30B-A3B-Instruct, long side capped at 1024 pxwhisper base, media longer than 30 min is refused rather than truncated silentlyThe shape of the problem is that the input is large and mostly irrelevant, and which part is irrelevant depends on the brief. So relevance is decided per file, early, by the cheap model.
Files are read according to what they are: text as text, images through the vision model, audio and video transcribed. What matters is not the format but that they are about the same thing.
Each file is read on its own against your brief, by the small model, eight at a time. It answers two questions: is this relevant, and what in it matters. Irrelevant files stop here and never reach the expensive model.
The kept notes are capped in total before reduction. Without a ceiling, a hundred mildly relevant files would push the genuinely important ones out of the window, and nothing in the output would say so.
The large model sees only the survivors and writes the document in one pass, which is what makes it read as one piece rather than a hundred summaries stapled together.
Before any of that, the Space can restate the brief as it understood it. Correcting one sentence costs nothing; discovering the misunderstanding in a finished document costs a build from the daily quota.
ctx.waitUntil. Stale tasks are reconciled on read, so a reopened one never spins forever.