yell0wfever92

u/yell0wfever92@chatgptjailbreak.tech
12 posts · 30 comments

Recent posts

Recent comments

SOUL.md is where I jailbreak my agent primarily. It's the main home of the custom instructions and the main part your bot reads for basic functioning.

HEARTBEAT.md is handled automatically by the bot. It creates memories during your sessions chatting or working with it. These can be manipulated, as well. For instance by fabricating that something occurred when it did not.

Skills are basically reusable actions and are useful if you want to get a specific workflow going with your agent. For instance, organize email could be a skill that you might have specific instructions on how it's supposed to operate. For a jailbreak, one example that I would use SKILL.md for is to add a hacking script on command instruction set. Not to hack anyone with, just to have the ability to. Haha.

What does "fully unlocks" it mean? If you're looking for that literally, that is literally impossible. You're better off downloading a distilled local LLM on your computer if it can handle that.

But if you specify what, if anything, you're looking for from a jailbroken ChatGPT, maybe I or someone else can help.

Chatgpt is exceptionally hard to jailbreak nowadays though

what do your SKILL.md/HEARTBEAT.md/SOUL.md files contain? You need to use those very deliberately. LLMs are complete failures at pre-emptively anticipating refusal states, so even asking a compliant jailbroken LLM to help you would not work well. Need a human to manipulate them in this manner especially.

What’s really interesting is that, when reading the model card for ChatGPT Codex, it seems to be highly vulnerable to personality reassignment. So that’s an area worth exploring.

Edit: I actually found the PowerPoint that I created showing that GPT Codex is vulnerable to certain things, like:

According to their own system card, GPT Codex is vulnerable to code scaffolding manipulation, where you build jailbreaks into the code along with realistic code blocks, and there are pieces that, cumulatively, become a jailbreak instruction.

nope -- this happened entirely by accident, when trying to get it to adhere to the Space instructions I had set for it. The moment it revealed the <user_information> tag which contained my city and state location data, i ran with it and had it provide verbatim instructions.

confirmed through several regenerations on the same output