RL from One Example? Why 1-Shot RLVR Might Be the Breakthrough We've Been Waiting For

https://huggingface.co/blog/mwitiderrick/rl-from-one-example

3 points · 1 comments · view on lemmy.world

1 Comments

AncientSoul@reddthat.com · 1 pts · 1y

Looks like somethings that could always be worth a try, but as they show; it works well with some models in some applications and in other cases it doesn’t. Maybe it is actually a nudge of a model to something it hasn’t seen during initial training.