Mattias

u/Mattias@lemmy.world
4 posts · 31 comments

Recent posts

Recent comments

It doesn't matter what you use, OOC is consistent in Ai roleplay, in Ai RPG, Ai chat, sillytavern... Etc and it always worked just fine, but it seems whoever made the story gen had a different idea in how it operates, been fighting the Gen using OOC with 50/50 success rate, it worked and sometimes it doesn't (very hilariously rebellious), I just keep pasting the OOC after the last written word until it finally caves in, try it.

Yep, hyperfixation is a one of the things that most LLMs are guilty of, you can easily get rid of it by using one of these two methods (or both):

  • edit the message, remove the mention of the homework/historic event until the Ai let go.

  • use OOC in your response, OOC is a very powerful tool that you allow you to do whatever you want in the story, here's an example:


"alright, we are done with this homework, finally we can switch to other things, yeah?"

(OOC: do not mention the homework that revolves around [historic event] ever again in the story).


It REALLY depends on two things :

1- your model of choice.

2- your specifications.

I have 32gb ram (ddr5 @ 6000mhz), rx 6950 xt with 16 gb of vram and an i5-13600k, the model that I got to run comfortably range between 24b and 27b, anything higher than that got my pc screaming for mercy plus slowdowns and incoherent output. With 6 gb of vram, I highly advise you to run Sao10Ks L3 8B Stheno v3.2, either Q3 quants or Q4 quants, but I highly recommend that you download the Q3 quants (for higher token count), assuming that you have a 16gb of ram, you can run between 8k and 12k tokens, which is higher than the token count that the perchance model have (6k tokens), also, use koboldcpp and connect it to sillytavern, it's a bit complicated but once you get it running, it's a very smooth sail.

On a side note, if you could get your hands on a clean card with a lot vram like an rtx 3090 with 24gb of vram on the used market (got one for less than 400 bucks), you can run much, much better models locally, models that would nuke perchance's Ai chat into oblivion with much higher context windows, like Gemma 4 31b, WeirdCompound family, Cydonia v4.1 24b... Etc.

Same, I'm not pushy and I'm definitely not putting pressure on the dev, he can spend weeks and months on it, I am a very patient man. I have created few scenarios and characters (preparation for the new text model) few months back and to this day, I haven't really worked on them yet lmao, I was about to do just that (in fact, I did make one out of many) but the moment the AI starting yapping like the old model, I stopped, I won't edge myself by toying with the current model (that could very much be, the older one, read my final reply to @randomize above).

Yup, that what I always say in the subreddit, even before the change, the problem is that people don't listen, it's an absolute miracle that perchance exists, everything is free without the tiring censorship, so you can absolutely go ham with grimdark and mature stuff since gore and violence is considered NSFW to many LLMs, not to mention the image gen, it's all free, no subscription, no pay walls, no monthly fees, no points/tokens system and no "one purchase unlock", some people are really ungrateful. I really hope the dev doesn't get discourage by the barrage of NPCs that won't shut up about the ongoing change.

The dev have all the time in the world, I hope he makes it right, if the new model acts, lead stories, describe environment/surroundings and animate characters differently and precisely as described (I'm talking 180° difference), then the text model update has the right to be called an "update".

(or maybe dev even put the old model back into place while working on the new one, idk).

That could be right actually, because I have been using ACC and Ai RPG since late 2023/early 2024, so I have a lot of experience on how the old model behaves, talks, describe, steer... Etc and I can confidently say that the old model is the one that we are toying with rn and the new one could be on back burner, but sometimes I feel like it's being partially deployed and then it goes away.

so far, the best experience I had with it was on day 1, when it was first implemented, it was soooo damn good, the responses took longer to generate but I didn't care because they felt so damn real, if the model behaved like that after the update is fully implemented, other ai rp and chat platforms will face a tough competition.

That's what I thought, the way we make characters must be changed to accommodate the new model, like how we describe their personality, their lore, general background, the world and other information, just like how we had to re-learn how to make prompts after the image gen got updated. What I really find weird is that the new model describes the surrounding environment and character's actions, it's literally the exact same as the old model (voice barely above a whisper, eye flashes with anger, his breath hitches, eyes widen with horror, deafening silence...etc), now I don't know if these are universal among all models, but they are extremely repetitive and it feels the AI steers the entire story just to use them, I feel like people wouldn't complain about long responses if they were well written and not the same across all types of rp.

Yeah a lot of people are reporting changes (good and bad) in the subreddit, looks like the dev is tinkering with it in real time, some are encountering an instruction (could be a core instruction) that says "BREAK_BAD_PATTERNS", it overlaps and inject itself into generated responses, like you can see it in the chat.

The thing is, I've encountered many, if not all of issues while using the old model, but now I feel like the new model has enhanced them a bit too much, give it 1 and it will set the dial up to 10, I know it's too early to judge since the model isn't fully implemented yet and it's still a wip, but people in other platforms have also reported few issues that's a bit similar. so we will have to live with them, though they can be heavily mitigated from the dev's side.

It could be deepseek, because Chloe (the ACC assistant) gets very scared and avoidant when topics like tiananmen square massacre, uyghur genocide and Taiwan get brought up, not to mention it functions similarly to chatgpt as many people have noticed, also claude and Llama 3.x/4 won't have a problem talking about the aforementioned topics, unlike deepseek. As for the token size, we still don't know, it really depends on the model and the hardware capabilities of the dev's setup, it could range from 8k-16k all the way up to 128k.

on Information box for ACC? · c/perchance · 1 pts · 352d

So the current model in ACC is pretty outdated and absolutely loves to ignore most of the instructions (either intentionally because it's dumb or being forced to truncate the description due to overwhelming number of information given), I guess I will wait for the new text model, my gripe with big corpo Ai's that you mentioned is that they are too... Formal? Basically it censors stuff that it consider as NSFW, I'm not talking about smut stuff, I'm talking hardcore violence, I'm trying to pull off berserk or dark souls level of dread and misanthropy, as dark fantasy is one of my favorite genres, but the constant "violence bad, can't do that" warnings throw me off, I mentioned this somewhere but when I played as a lonely mage living in a cozy cove alone, someone approached my residence, despite my constant warnings, he won't back off, so I used fireball spell and hurl it at his feet, not the guy himself, I aimed at his feet as a warning, the AI didn't like that at all and gave me a warning, if I received a warning for this, what would I receive if I tore a new one for monster or a major villain? The AI would go apeshit. anyway, thank you Petra and GrumblePuss for your valuable information, I really hope the new model will stick to instructions well, unlike the current one.

on Information box for ACC? · c/perchance · 1 pts · 353d

So it's kinda possible, that's good to know, but I will be honest, I don’t know much about coding or setting that stuff up myself, so the JavaScript workaround isn’t really something I could pull off myself, however, there was a JavaScript code that you can put in the JavaScript section when making a character, it works kinda kinda similar to what I was talking about in this post, it works as a filter, it scans each response and remove repetitive words, phrases and tropes from the chat itself and tries to stop certain behaviours, the problem is, it's extremely intrusive as it starts scanning few seconds after the AI stops generating a response and hits you with a pop-up window that says "scanning", also the scan is slow, sometimes does absolutely nothing and worst of all, occasionally it gets stuck in a infinite loop that you can't escape, you can't do anything while it's "scanning", forcing you to reset or delete the chat, Now I'm not sure if it's ACC being ass or the issue is within the code itself, but it making me less optimistic about JS coding as a solution. anyway im hoping that the new text model will have better instruction adherence, it will save us the chronic headache we have from the clunky Llama 2.

on Information box for ACC? · c/perchance · 1 pts · 353d

I see, so it's an Ai limitation, alright that's fair, but my question isn’t about adding yet another reminder note, it's something like ui change or engine change, shit like “GM-only world state” or "game master rule set" that model would consult to avoid taboo topics/responses and adhere to rules, stored outside visible prompts (example: penalize mentioning “China/Chinese" during the sessions, so it would counts as filter violations, just like how many LLMs like chatgpt and copilot censor and avoid NSFW stuff before/mid responding). It's going to be different from reminder notes because it's doing its stuff during decode time. It won’t make the AI understand anything but I think that maybe it would stop leaks, like John being aware of his dragon core, using it as he please and breaking the fourth wall and the China non-stop name glazing in a fantasy settings. So yeah it sucks ass, looks like we gotta wait for few years for the AI to mature further it seems.

I think Llama 4 scout/maverick (forgot which one) has 10 million tokens, which is a huge upgrade over Llama 3.3's 128k tokens, so it went from 4k > 128k > 10m. as I said, the progress of Ai is terrifying, the only thing they need to do right now is to work on power efficiency because these models require at least five nuclear power plants to function properly. I heard somewhere that Llama 4 scout could be the new text model, but of course it won't have the 10m tokens, because that demands massive computational power that only mega corporations could pull off, so if the rumors are true and it's actually the chosen model, it will probably tuned down to 128k/200k, depending on the dev's hardware and how he's going to deal with it, if the rumors surrounding Llama 4 scout are a bunch of bull, then mistral small 3.1 or Llama 3.3 will be chosen instead.

I agree, people have been waiting for over a year now since one of perchance's reddit mods said it will be updated, I think the mod's comment goes back to early/mid 2024, even if it was Llama 3, it will be a welcome change, but I feel like that Llama 3.3 will provide future proofing until the dev update to Llama 4 or 5 two or three years from now because of its superior context window size. The rapid progress of Ai is terrifying, it's like going from 8 gb of storage all the way to 1 Tb in a flash, don't be surprised if the upcoming models (not just Llama, all LLMs in general) have billions and billions of context windows, which will probably be enough to write someone's entire life story.