What strikes me most is the size of the model. 2.6B is not that much if we think about what it is being asked to do across tool use and multi-step workflows, especially on-device. At the same time I wonder how reliable this model remains once it moves outside the harnesses and tasks used during training. That is where we really understand whether we are looking at something that is actually usable.
Very interesting. We talk a lot about models, but much less about dataset composition. Small shifts in data distribution can easily have a greater impact than architectural improvements.