Distilling step-by-step: Outperforming larger language models with less training data and smaller model sizes

https://blog.research.google/2023/09/distilling-step-by-step-outperforming.html

28 points · 3 comments · view on lemmy.world

3 Comments

noneabove1182@sh.itjust.works · 3 pts · 2y (2 replies)

Woah this is pretty interesting stuff, I wonder how practical it is to do, I don't see a repo offering a script or anything so may be quite involved but looks promising. Anything to reduce size while maintaining performance is huge at this time

Zetaphor@zemmy.cc · 2 pts · 2y (1 reply)
noneabove1182@sh.itjust.works · 1 pts · 2y

Somehow this is even more confusing because that code hasn't been touched in 3 months, maybe just took them that long to validate? Will have to read through it, thanks!