Orthrus-Qwen3: up to 7.8×tokens/forward on Qwen3, identical output distribution

https://github.com/chiennv2000/orthrus

Crossposted from https://lemmy.ml/post/47429470

Paper: https://arxiv.org/abs/2605.12825

10 points · 2 comments · view on lemmy.world

2 Comments

BB84@mander.xyz · 3 pts · 103d

My oversimplified and possibly wrong understanding: this is like speculative decoding, but instead of a separate draft model (which does its own prompt processing), they use some diffusion thing strapped on top of the main model. The diffusion reuses the high-quality prompt processing result of the main model.

The 7.8x faster claim sounds almost too good to be true. But even if we get like 3x then this is still a huge revolution in localLLMing.

tristynalxander@mander.xyz · 2 pts · 103d (2 replies)
[ removed ]
BB84@mander.xyz · 2 pts · 103d (1 reply)

They said they're working on Orthus for Qwen 3.5. It'll be amazing!

tristynalxander@mander.xyz · 6 pts · 103d
[ removed ]