A chat model trained from scratch on a Mac, now running in your browser.
tinyurl.com/2y8xft2k
Scan now and follow along: the model runs on your own phone or laptop. The question this whole
project asks is in the title.
The idea · 2026-10-03
Train it from scratch, then make it ternary
Not a big model shrunk: a small one grown
Nearly every weight is −1, 0 or +1
Trained on one Mac, learning from a stream
“MiniAGI + QAT”: the design, in two words
Started from an open small-model design (mini-AGI), deliberately derived from it, and trained
it ternary (quantisation-aware training) from step 14,655 on. The first milestone was just to reproduce the original exactly, and it matched to
six decimal places before anything changed.
What was built · as of 2026-10-07
Bytes in, one looping block, bytes out
No tokenizer: raw bytes
32 experts, 8 per byte
Thinks longer on harder bytes
Byte-level means it can read anything but has to learn spelling itself. A mixture of experts
keeps each step cheap. The looped block with adaptive halting lets the model decide how many passes each byte
gets, from 1 to 24.
Make the number bigger · 2026-10-04 to 10-06
Grow until it stops paying
109M → 310M active in a day
Context 4K → 8K
Memory and swap set the limit
The mandate was to keep every setting going up until a measurement said stop: worse
usefulness, the GPU can't hold it, or it slows down. It hit memory first; 256 experts were merged to 128.
Grey is total weights, green is the part active for each byte.
What went wrong: rigor · 2026-10-04 to 10-07
Every failure became a mechanism
Test data leaked into training again and again, in ways that got subtler each time: whole
files, then shared openings, then reworded copies. Each time the fix was structural, not a promise: the held-out
sets now live on a different machine, and every dataset is swept for 8-word overlaps before anything trains on it.
What went wrong: autonomy · 2026-10-05 to 10-07
Make it run while nobody watches
The goal moved to hands-off: programs do the routine work, an alert wakes the lead agent, and
every alert is also a reason to make the system sturdier. Programs once changed training settings on their own,
so the next piece is a single owner for settings, the regulator, which is approved but not built yet.
How it runs · 2026-10-07
Programs do the routine; people decide
While it runs: the trainer reports to the supervisor, which restarts it, steps it down when
memory runs short, follows new data versions, and changes only the settings the owner made automatic.
Checkpoints go to the desktop to be scored against data the trainer never sees; scores feed the rules and
the site. Anything that could stop training or lose data raises an alert that wakes the lead agent, which
fixes the cause. Architecture and anything leaving the machine stay the person's call.
Change of course · 2026-10-07
Smaller, shorter, and only chat
Cut from the big run: 109M weights, 32 experts
2,048-byte context: short chats
Open chat data, plus chats written for it
The big run was paused, kept exactly, and a smaller copy cut from it to push toward coherence.
Most data is open chat datasets (conversations under 2K bytes); 15% is synthetic chats written to show the
character we want.
The charter · 2026-10-07
Teach intent, not rules
Two open models act out chats: a person, and Nirca Mini
The charter says who Nirca Mini is, and why
Honesty first: “I’m not sure” beats a confident guess
The charter steers the synthetic chats; the page sends no system prompt. It went through six
versions in one day. The turn was from lists of rules to reasons with examples: a model learns the intent from
examples better than a rule list. The full text is in the Charter tab.
Where it stood · 2026-10-07, honestly
It had learned the shape of a chat, not yet the sense
Graded coherence: about 0.01 out of 1
Chat loss fell on the big run, rose on the Mini
Read as too many repeats of too few synthetic chats: cut to 15%
An open model grades 148 replies for being sensible and specific. Every big-run checkpoint
scored zero; the Mini scores about one in a hundred. Validation loss on chat rose on the Mini (scored at 2K
context, so not directly comparable), read as overfitting the small synthetic set. About 1.3 billion bytes seen
so far.
Live demo
The same model, on your GPU
33 MB, then cached
WebGPU: Chrome, Edge, Safari
Same replies as the reference code
Switch to the chat. It loads on open. Try "hi" and "What is the capital of France?". Nothing
leaves the page. If the demo stalls, the point still stands: the whole model is a 33 MB file.
What's next
Coherence first
One owner for training settings: the regulator
Fewer repeats, more multi-turn chat
Newer Nirca Mini weights, public, loaded by this page
tinyurl.com/2y8xft2k
Close on the link. The model in the browser today is from its first day of training; newer
checkpoints will follow, and the page will offer them.