Dense and Moe models under 36B, please

#31
by Duonglv - opened

The community has high expectations for small dense or MoE models—those under 36B parameters.

I see that almost all Chinese AI providers, such as DeepSeek, Qwen, Z.ai, …, no longer focus much on truly small models. They release medium-sized models in the range of 120B to 360B parameters. These models are good for companies, but not for individuals or startups with very limited resources.

If this trend remains, we can only expect Google to release the next Gemma models.

Sign up or log in to comment