Post
814
20K parameters can tell a story. 🚀
🤗 We trained a ~20k-parameter Transformer that can actually write stories!
raincandy-u/MacroStories
→ ~50× smaller than the 1M-parameter TinyStories model
→ ~3,000× smaller than AlexNet
→ 81 KB in FP32
yayyy the whole model. ૮ ˶ᵔ ᵕ ᵔ˶ ა
She has a 32-dimensional hidden state, a 378-token vocabulary, and just one decoder block — recurrently applied 4 times with shared weights.
Despite having only 19,969 parameters, she can maintain a simple narrative across 100–300 words: establish a goal, encounter a problem, take relevant actions, and reach an outcome.
She runs extremely fast on CPU — no GPU required. The entire model is tiny enough to load almost instantly! ☺️
🤗 We trained a ~20k-parameter Transformer that can actually write stories!
raincandy-u/MacroStories
→ ~50× smaller than the 1M-parameter TinyStories model
→ ~3,000× smaller than AlexNet
→ 81 KB in FP32
yayyy the whole model. ૮ ˶ᵔ ᵕ ᵔ˶ ა
She has a 32-dimensional hidden state, a 378-token vocabulary, and just one decoder block — recurrently applied 4 times with shared weights.
Despite having only 19,969 parameters, she can maintain a simple narrative across 100–300 words: establish a goal, encounter a problem, take relevant actions, and reach an outcome.
She runs extremely fast on CPU — no GPU required. The entire model is tiny enough to load almost instantly! ☺️