Bytedance is training an AI model with up to ten trillion parameters: That's three times the size of Moonshot's Kimi K3, currently the largest Chinese model.
https://the-decoder.com/chinas-largest-ai-model-is-being-developed-at-bytedance/
https://the-decoder.com/chinas-largest-ai-model-is-being-developed-at-bytedance/
6 Comments
RickyRigatoni@piefed.zip · 36 pts · 30d
yeah bro just ten trillion more parameters man all we need is more parameters and ai will be worth the investment bro just give us 300 more datacenters to hold all these parameters and make the homeowners pay our power bill because we're still not profitable after ten trillion parameters please bro i'm begging you
DrakeAlbrecht@lemmy.world · 8 pts · 30d
eicker@lemmy.world · 14 pts · 30d
ByteDance building China’s largest model while running TikTok is fascinating: few companies have that combination of compute, money and an absurdly large stream of real world human behavior. The AI race isn’t just US labs versus China anymore: it’s ecosystems versus ecosystems.
Squizzy@lemmy.world · 5 pts · 30d
Meta are losing their AI market
Armand1@lemmy.world · 7 pts · 29d
Having run models locally, RAM use seems to be almost directly proportional to number of parameters. 8 Billion parameters requires approx 8GB of VRAM at 1/4 precision.
Therefore, if this pattern holds you somehow need 10 Terabytes of VRAM at 4K and 40 Terabytes at full precision.
I think I saw some estimates that Claude's Opus models may be and Opus model equivalents may be at around 100B parameters (100-400GB VRAM).
TLDR its clear why RAM is so expensive.
eager_eagle@lemmy.world · 3 pts · 29d
for inference you're only counting active parameters towards VRAM, and some labs / models don't train at 32b precision, or even use the same precision for different parts of the network