NNL architecture notes

Four tools.
Mostly the same machinery.

cloudflare workers · d1 · durable objects · r2 · hugging face spaces

Speech, vision, language and retrieval are the same four pieces every time. What differs between the four systems is the arrangement, which is why a new one takes weeks rather than a restart. This page is what is actually running.

01 Topology

Two deployments, one database, four Spaces. The browser never talks to a model directly except on IRIS, which is deliberate and explained below.

Marketing sitennlabs.pl, a Worker serving static assets
Applicationapp.nnlabs.pl, Cloudflare Pages in advanced mode, a single _worker.js at the edge
DatabaseD1, one instance shared by auth, chat history, LUNA tasks and ARIA events
Realtimea standalone Worker holding the ARIA Durable Object and R2, reached by service binding
Modelsfour Gradio Spaces on Hugging Face, driven server-side over the queue API
SessionsPBKDF2 at 100 000 iterations, HttpOnly cookie, 30 days

02 How a request survives

Most of the engineering here is not in the models. It is in the space between a browser that can close at any moment and a Space that takes minutes to answer.

Page authorisation at the edge

Tool pages are checked against the session cookie before the HTML is served. A client-side redirect is not authorisation; it is a suggestion that view-source ignores.

protected: /nina /iris /luna /aria

Signed use tokens

The Spaces are public, so the model would be free to anyone who found the URL. Each use mints an HMAC-signed token with a short life, and the Space refuses anything else.

HMAC-SHA256 · TTL 120 s · unique id per token

Turns outlive the tab

NINA and LUNA runs are driven from the Worker with ctx.waitUntil and written to the database as they complete, so closing the browser loses the stream but not the answer.

NDJSON progress to the client · partial answers persisted

One active session per account

A claim plus a heartbeat. The newest tab wins and the older one is told, rather than both quietly spending the same budget.

heartbeat 20 s · stale after 45 s

Quota on delivery

A use is counted when a model returns something. Where a token has to be minted before the Space will start, the use is taken up front and refunded if nothing came back, guarded by the token id so one failure cannot be claimed twice.

10 uses / 5 h · LUNA 3 / day · ARIA 20 / 5 h

03 Decisions worth defending

Open checkpoints throughoutWhisper, Qwen, CLIP, pyannote, gemma. Nothing in the chain is a wrapper around one vendor's endpoint, so nothing in the chain breaks when that vendor changes its mind.
Server-driven model callsThree of the four are driven from the Worker rather than the browser. The client cannot be trusted with the quota, and the answer should not depend on the tab staying open.
IRIS is the exceptionIt is called directly from the browser and is the odd one out architecturally: a different Gradio generation, a raw API surface, free-tier hardware. It works, and it is the next thing to bring in line.
Media is never storedUploads are processed and dropped. Derived text is kept, media is not, except ARIA attachments which have to stay visible in the conversation.
Two peopleEvery decision above is also a decision about what two people can maintain. That constraint did more shaping than any of the others.