LUNA Latent Understanding & Narrative Assembly

Many files in.
One document out.

gradio space · api_name: build / clarify · map-reduce

A hundred mixed files do not fit in any context window, so LUNA does not try. It reads each one separately with a cheap model, keeps only what the brief needs, and hands the survivors to an expensive one.

01 Stack

Two language models with different jobs and very different price tags, plus a vision model for anything that arrives as a picture.

MapQwen/Qwen2.5-7B-Instruct, temperature 0.0, 900 tokens per file, 8 workers in parallel
ReduceQwen/Qwen2.5-72B-Instruct, temperature 0.15, up to 4000 tokens
ImagesQwen/Qwen3-VL-30B-A3B-Instruct, long side capped at 1024 px
Videoup to 8 frames at 768 px, audio track transcribed separately
Audiowhisper base, media longer than 30 min is refused rather than truncated silently
Timeout180 s per model call, 8 min hard cap on one build

02 Map, then reduce

The shape of the problem is that the input is large and mostly irrelevant, and which part is irrelevant depends on the brief. So relevance is decided per file, early, by the cheap model.

Intake

Files are read according to what they are: text as text, images through the vision model, audio and video transcribed. What matters is not the format but that they are about the same thing.

MAX_FILES 100 · 25 MB per file · 200 MB per media file · 500 MB per task

Map

Each file is read on its own against your brief, by the small model, eight at a time. It answers two questions: is this relevant, and what in it matters. Irrelevant files stop here and never reach the expensive model.

PER_FILE_CHARS 16000 in → MAP_NOTE_CHARS 4000 kept

Budget

The kept notes are capped in total before reduction. Without a ceiling, a hundred mildly relevant files would push the genuinely important ones out of the window, and nothing in the output would say so.

REDUCE_BUDGET 60000 chars

Reduce

The large model sees only the survivors and writes the document in one pass, which is what makes it read as one piece rather than a hundred summaries stapled together.

REDUCE_MAX_TOKENS 4000 · TEMPERATURE 0.15

Clarify

Before any of that, the Space can restate the brief as it understood it. Correcting one sentence costs nothing; discovering the misunderstanding in a finished document costs a build from the daily quota.

api_name clarify · no quota charged

03 Decisions worth defending

Two models, not oneThe same model doing both jobs is either too expensive per file or too weak at the final write-up. Splitting them makes each choice independently defensible.
Temperature 0.0 on mapExtraction should be reproducible. Two runs over the same file should keep the same facts.
Explicit character budgetsEvery stage has a stated ceiling. Silent truncation is the failure mode that produces a confident document with a hole in the middle.
Quota charged on deliveryThe daily count moves when a Note is stored, not when a build starts. A paused Space or a model error costs the user nothing.
Build survives the tabThe Worker drives the Space through ctx.waitUntil. Stale tasks are reconciled on read, so a reopened one never spins forever.