New open-weight 🐋 DeepSeek V3. 685B MoE. Beats Claude 3.5 Sonnet on Aider coding benchmark

https://huggingface.co/deepseek-ai/DeepSeek-V3

Absolutely humongous model. Mixture of 256 experts with 8 activated each time.

Aider leaderboard: The only model above 🐋 v3 here is OpenAI o1. DeepSeek is known to make amazing models and Aider rotates their benchmark over time, so it is unlikely that this is a train-on-benchmark situation.

Some more benchmarks: on Reddit.

36 points · 1 comments · view on lemmy.world

1 Comments

xodoh74984@lemmy.world · 5 pts · 1y (1 reply)
[ removed ]
BB84@mander.xyz · 3 pts · 1y

Someone managed to run it on a cluster of Mac Minis lol https://blog.exolabs.net/day-2/

toothbrush@lemmy.blahaj.zone · 1 pts · 1y
[ removed ]