flasgship pretrained models
🤏 quecto mode
appvoid
appvoid
AI & ML interests
singularity through byte-level tokens and small language models
Recent Activity
posted an update about 5 hours ago
- GLM 5.2
- Flux 3
- New Qwen model
- New small model leaderboards
- Lots of people finetuning smol models.
- Some even under 12 year olds clauders are here (was not on my bingo card this year)
- ChatGPT's Sol became a lot faster this week
- LFM2.5 2.6b
- Kimi K3 (though only a few will run it)
- New Ling 3.0 Tiny
- New video model that is making south park videos?
- Deepseek v4 flash being more honest than bigger models
- The new model from meta
Everything Everywhere All At Once liked a model about 6 hours ago
BananaMind/BananaMind-2-Micro repliedto Banaxi-Tech's post about 6 hours ago
We're excited to release BananaMind 2 Micro, our smallest model yet.
It fits a compact architecture in only 2.9M parameters achieving the highest parameter efficiency on BananaMind Base Bench against comparable models.
It achieves comparable performance to GPT S2 5M and GPT S 5M at almost half the size while beating CMA 1M Mini.
BananaMind 2 Micro achieved the #1 spot on the Open SLM Leaderboard for the sub 3M category (not added yet but it achieves #1)
For the training we used Muon + the XSA Refresh Gate with a 5e-2 lr for Muon and 4e-3 for the 1D weights.
Its score on our efficiency measure is 0.326 getting the first place with Syn 2.6M on the second place scoring 0.291 and GPT S 5M at 0.235*
Check it out at https://huggingface.co/BananaMind/BananaMind-2-Micro and follow us at:
@vovaRL
@DedeProGames
@Banaxi-Tech
https://huggingface.co/BananaMind
Our new releases aren't stopping 🚀 August 13-14 BananaMind 2 Pro