Dean Byrne PRO
AI & ML interests
Recent Activity
Organizations
I built an arena where tiny decoder-only LMs (50K–250M params) play Tetris zero-shot. There is no fine-tuning and no game data. They only use what they picked up from pre-training on text.
How it works:
- For every piece, the engine simulates each legal placement and describes the result in plain English ("clears one line, creates no new holes, keeps the stack low…").
- The model never sees the grid. It reads each description, and the arena compares log P(" good move") with log P(" bad move"). The best-rated placement is played.
- Every player gets the same piece sequence, so it's a fair race.
- There are two protocols: Guided (the rules are in the prompt) and Blind (no rules, only pre-training knowledge).
Two ways to play:
- Match: pick any models (even your own, custom architectures welcome) and watch them play side by side on retro 8-bit boards.
- Ranked: press Play and the arena picks up to 4 models at random from a curated pool of 29. Nobody chooses their opponents, so Elo can't be farmed. Matches run on the server and count even if you close the tab.
First results (~225 ranked matches):
- gpt2 (124M) leads with 1283 Elo, but SupraNeo-4M (4M) is right behind at 1239. Next come LowOnMind-5M and BananaMind-2.1-Pico (1.5M!).
- Model size barely predicts Elo (r ≈ 0.06). Survival does (r ≈ 0.9): the models that avoid holes and keep the stack low are the ones that win.
Every ranked match (seed, model commit SHAs, scores, Elo before/after) is logged in a public dataset.
▶ Play: DedeProGames/SLM-Tetris-Arena
📊 Results: DedeProGames/lm-tetris-arena-results
Want your model in the Ranked pool? Drop it in the comments!
Good catch. Shipped decision_config.json was stale vs the weights.
DecisionAgent does not read that file. It reads temperatures from the checkpoint’s embedded decision_cfg. On-Fly-Jev.pt embeds [1.0, 1.0, 1.0] and has no per-bucket temps. So 0.12 never ran. The clamp never ran. Every number I reported is T = 1.0.
Accuracy does not care. Temp does not move argmax, and it does not move the noul 0.5 cut.
You were right about choice:2. That fit was junk. Hold-out was 12 items: 8 SST2, 2 IMDB, 2 other. It slammed into the 0.1 floor. SST2 got worse, not better:
• T = 1.0, choice-form ECE 0.254
• T = 0.5, 0.297
• T = 0.12, 0.310
One correction: about half of SST2 and IMDB are noul:2, not choice:2.
What I changed:
• Config now matches the weights.
• A bucket needs at least 100 items or it falls back to the question-type temperature.
• Fits clamp to [0.5, 5.0], same as the agent.
I refit on 4,093 held-out benchmark items, kept off the eval set. Ceiling is real. At T = 1, typed is fine (ECE 0.045) and the benchmarks are overconfident (0.197). Fitted temps pull benchmarks to 0.085 and wreck typed (0.245). One T per bucket cannot do both jobs.
Published model stays T = 1. Benchmark-calibrated copy is in calibrated/. README says both now.
It's a 4B model that turns spoken or typed English into actions in Unity scenes. Say "put the red mug on the table" or "turn on the lamp", and it returns the tool call your app executes. If a command could mean two objects, it asks which one. If it can't do something, it says so instead of guessing.
Everything runs on the user's machine through llama.cpp: no API key, no internet connection. The Q4_K_M GGUF is 2.5 GB and needs about 3 GB of GPU memory, so it fits on a 4 GB laptop GPU and usually answers in one to three seconds. It is fine-tuned from Qwen3-4B with QLoRA on about 20,000 English conversations.
Results:
- 83.8% on 499 human-written ALFRED instructions (right action on the right object). The base model, Qwen3-4B, scores 57.1%. The strongest of the five other models we tested, from 1.7B to 120B parameters, was Ministral 3 14B at 67.1%.
- 91.7% on object types it never saw in training.
- 97.3% on 440 commands run through a live Unity scene.
There is also a Unity package that starts the model, describes the scene to it and carries out its tool calls. You install it from the Package Manager with a Git URL.
Model: ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF
Unity package: ErenAta00/Maverick-Unity
Full write-up: https://huggingface.co/blog/ErenAta00/maverick-4b-unity-xr-agent
Built at the Extended Reality Laboratory (XRLab), Manisa Celal Bayar University:
Released under Apache-2.0. Feedback and bug reports are welcome in the Community tab.
Second Jev-style model. First was Byrne-Jev (70M SpikeWhale). This one is a 96M spiking trunk - every unit is a copy of one of 100 real MaleCNS fly neurons - plus a typed-decision head. One forward pass. No generated text.
I built it for speed.
Typed-decisions test split, 400 cases, 2,000 decisions, 5 questions per case, local GPU:
• 7.3 ms p50 per decision (~137 / s)
• 36.6 ms p50 / 53.5 ms p95 per case of 5
• ~27 cases / s
Same protocol vs the others:
• On-Fly-Jev: 36.6 ms / case, ~137 decisions / s
• Byrne-Jev: 110.7 ms / case, ~45 / s
• ModernBERT-base: 349 ms / case, ~14 / s
• TypeSafe Jev 1.13 (hosted, so network is in it): 710 ms / case, ~7 / s
About 3x Byrne-Jev, 9.5x ModernBERT, 19x TypeSafe Jev per case.
Live ViZDoom, 1 question per tick including game I/O: 23.4 ms (~43 / s).
• Accuracy 0.666 (Byrne-Jev 0.630, Jev 1.13 0.727)
• ECE 0.045, same as Byrne-Jev, about a third of Jev 1.13
50/50 merge of two checkpoints from one run. Research artifact, not a chatbot. More videos are on the card.
Check it out and like it!
AxiomicLabs/Tiny_Theory_of_Mind
I'm officially canceling my Hugging Face Pro subscription today.
I supported this platform because it stood for true openness and neutrality. This acquisition by NVIDIA fundamentally changes that.
Here’s why I’m against this deal:
- Neutrality is dead. NVIDIA is a US-based company. This means US regulations will inevitably dictate platform policies, creating direct pressure on Chinese developers and anyone building open-weight models outside the US.
- Community over bureaucracy. NVIDIA is a massive, slow-moving corporation. This acquisition will likely drown the community in corporate processes and commercial interests. Soon, uploading a simple finetune might become a bureaucratic nightmare.
- Open vs. Proprietary. Hugging Face was built on open-source ideals. NVIDIA? They are a fiercely proprietary hardware company with a minimal track record of meaningful open-source contributions. They sell chips, not freedom.
- And to add insult to injury, NVIDIA has practically abandoned consumer RTX GPUs in 2026 to chase data center profits. Why would I pay them for "openness" when they've turned their back on the very developers who built this ecosystem?
I paid for openness. Not for a corporate takeover.
🤗 was about community.
ForgeWorks/ForgePlex-M1-6M just dropped from
Achieving an Intelligence Index of 6.87 and taking #22 in the <10m category on the AxiomicLabs/Open_SLM_Leaderboard, very impressive work for a first model.
Give it some love!