NNL architecture notes
Speech, vision, language and retrieval are the same four pieces every time. What differs between the four systems is the arrangement, which is why a new one takes weeks rather than a restart. This page is what is actually running.
Two deployments, one database, four Spaces. The browser never talks to a model directly except on IRIS, which is deliberate and explained below.
nnlabs.pl, a Worker serving static assetsapp.nnlabs.pl, Cloudflare Pages in advanced mode, a single _worker.js at the edgeMost of the engineering here is not in the models. It is in the space between a browser that can close at any moment and a Space that takes minutes to answer.
Tool pages are checked against the session cookie before the HTML is served. A client-side redirect is not authorisation; it is a suggestion that view-source ignores.
The Spaces are public, so the model would be free to anyone who found the URL. Each use mints an HMAC-signed token with a short life, and the Space refuses anything else.
NINA and LUNA runs are driven from the Worker with ctx.waitUntil and written to the database as they complete, so closing the browser loses the stream but not the answer.
A claim plus a heartbeat. The newest tab wins and the older one is told, rather than both quietly spending the same budget.
A use is counted when a model returns something. Where a token has to be minted before the Space will start, the use is taken up front and refunded if nothing came back, guarded by the token id so one failure cannot be claimed twice.