c/ai_ · by cm0002@lemmy.world · 1yRL from One Example? Why 1-Shot RLVR Might Be the Breakthrough We've Been Waiting For https://huggingface.co/blog/mwitiderrick/rl-from-one-example3 points · 1 comments · view on lemmy.world
1 Comments
AncientSoul@reddthat.com · 1 pts · 1y
Looks like somethings that could always be worth a try, but as they show; it works well with some models in some applications and in other cases it doesn’t. Maybe it is actually a nudge of a model to something it hasn’t seen during initial training.