A LLM benchmark that gave a hard programming tests to gpt 5.6, but for much more languages than common benchmarks

https://danluu.com/pl-tokens/

He's right that python/js tend to score better on most tasks. OpenAI is very bad at less popular languages. Especially J, so that is a very bad skew in his results.

-2 points · 0 comments · view on lemmy.world

0 Comments

No comments yet.