a post to donate types of prompts that could be used as a benchmark.
criteria guideline:
- problems that llm had struggled to solve. ( and somewhat desirable )
- valuable usecases
- centric to various other problems ( with value )
- located in a domain that underperforms on all tasks related to it
* soft push for actual personally useful problems. not toy demonstrations/tests .
* no criticizing/debating . this is write-only
* soft focus for smaller models ( rationale being: no really hard problems that are cutting-edge math,science,engineering ) . eg ok to be using large models but for not cutting-edge problems that a smaller model definitely could never handle.
* one post per user
* vague description of problem/prompts is ok.
1 Comments
leanleft@lemmy.ml · 1 pts · 13d