I mean most SLM labs are doing the same thing though, go look at BananaMind, Axiomic Labs, and FromZiro. So i wasnt doing anything out of the normal, and its the same training data as Pebble-25M and Pebble-10M so it isnt 100% the dataset.
There are still some interesting improvements in these models: - Compatible with non-CUDA devices - Vocabulary increased to 16K tokens - Context length increased to 16K tokens
For now, there won't be any more Pebble releases for a while. We're going to take some time to experiment with other approaches and hopefully make the next generation a monumental leap over this one.
There are still some interesting improvements in these models: - Compatible with non-CUDA devices - Vocabulary increased to 16K tokens - Context length increased to 16K tokens
For now, there won't be any more Pebble releases for a while. We're going to take some time to experiment with other approaches and hopefully make the next generation a monumental leap over this one.
Thank you for that link, i might train future models with mamba3 right now though im going to stick with mamba2, mamba2 vs mamba3 will probably be one of the things i test in the hunt for the optimal mamba models.