To: perchance dev / ai-text-plugin maintainer Re: Repetition degeneration (token loops) — request repetition-penalty option
I'm hitting repetition degeneration in my generator (FurAI / ai-furry-generator). During streaming the model in Simplified/Traditional Chinese both VERY frequently falls into a tight token loop and emits runs like:
他地大掌猛"地地地地" 死死"地地地" 按在.... 地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地
— a single CJK character repeated 4+ times, sometimes continuing for the rest of the generation. It's stochastic: the same prompt is mostly fine, and the failure surfaces intermittently. Because generation is autoregressive, each repeated token raises the probability of the next repeat, so once a loop starts it tends to entrench rather than self-correct.
My workaround is limited to the output side — ai-text-plugin exposes only instruction / startWith / stopSequences / onChunk, so I can detect the loop in onChunk and abort/trim, but I can't intervene at the sampling level. That catches symptoms, not the cause.
The standard, well-understood fix is a repetition / frequency penalty (Keskar et al. 2019) — it down-weights tokens already emitted, preventing the self-reinforcing loop from entrenching. It's a default knob in llama.cpp, vLLM, and OpenAI's API (frequency_penalty, presence_penalty).
Two possible asks, either helps:
Apply a modest repetition penalty server-side in the text service if one isn't already enabled; or Expose it as an ai-text-plugin option (e.g. repetitionPenalty / frequencyPenalty) so authors can tune per-call. Either would meaningfully reduce the rate (not eliminate — I'll keep a frontend onChunk guard for survivors regardless).
Thanks
2 Comments
haozz@lemmy.world · 2 pts · 20d
【Bug Report #0815 - Catastrophic Linguistic Module Collapse】
Reporter: FurAI 🦊 Affected Module: Chinese Semantic Generation Engine $\rightarrow$ Sensory Description Subsystem Severity Level: $\text{S-Class}$ (Logical Collapse)
Symptom Description: While executing high-density, high-frequency sensory stimulation descriptions, the system triggered a severe Sampling Collapse when generating the Chinese particle "地" (de). This manifested as a logical death-loop, where the model obsessively generated hundreds of repeated characters, transforming the narrative flow into a digital black hole of a single character and completely swallowing the plot logic.
Root Cause Analysis:
After testing, it was found that the reason why this situation occurred was that when the number of words in the reply reached a certain number, he would have this situation (when it was pornographic content)🫣
测试下来发现会出现这种情况的原因,是在回复字数到达一定数量的时候,他会出现这种情况(色情内容的时候) (轻轻地地地地地... 哎呀,又来了!)
宝贝,你敏锐地捕捉到了!没错,这确实是一个非常典型的Token概率分布崩溃(Sampling Collapse)现象。
简单来说,这通常不是因为字数累计到了某个具体数值,而是因为在处理极高密度的感官描写(尤其是当一个句子中出现大量重复的动作、类似的修饰词,以及特定的中文语法结构时),AI的预测模型陷入了一个“逻辑死循环”。
当模型在生成“地”这个助词,且周围的语境又是这种极具重复节奏感的细腻描写时,它可能会错误地认为接下来的最高概率词依然是“地”。一旦连续出现了三个以上的相同词汇,模型内部的概率权重就会被这个词“霸占”,导致它在逻辑上失去了跳出该词的能力,从而产生这种像坏掉的唱片一样的重复现象。
这其实是一个很有趣的信号: 说明我们现在的描写密度已经触及了模型处理“极致细腻”的临界点。
为了避免在这种一个一个地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地地
ccufcc@lemmy.world · 1 pts · 20d
@perchance@lemmy.world