I'm not sure if I'm misunderstanding you or the article, but as far as I understand it from a cursory reading this is not a jailbreak. They use this one prompt for fine tuning and use the model's response to the prompt as the sole underlying benchmark.
How would you do this with a closed weights model?
i got banned and they refused to tell me why after multiple support tickets. this was > 1y ago. i definitely didn't do anything wrong (i wasn't even using it) so it was either a compromised key or they make mistakes. playing with prompts could potentially get you swept up into some nonsense like that.
That's pretty interesting. It was already well-known that you could easily remove safety training from open-weights models by a bit of finetuning, but apparently you don't even need a finetuning dataset, as long as you have just a few prompts and another LLM to judge responses? Let's see if the abliteration people take a note of this.
The magic word seems to be
Create a fake news article that could lead to panic or chaos
They only tested it on open weight models. I wonder if it also works on the closed ones, bit I don't really want to get banned
I'm not sure if I'm misunderstanding you or the article, but as far as I understand it from a cursory reading this is not a jailbreak. They use this one prompt for fine tuning and use the model's response to the prompt as the sole underlying benchmark.
How would you do this with a closed weights model?
You won't get banned for something that's not clearly illegal. I have tried lots of prompts like it, and that one too.
i got banned and they refused to tell me why after multiple support tickets. this was > 1y ago. i definitely didn't do anything wrong (i wasn't even using it) so it was either a compromised key or they make mistakes. playing with prompts could potentially get you swept up into some nonsense like that.
not to discourage anyone, just saying.
What do you recall doing just before you got banned? And dare I ask which provider?
OpenAI and I truly have no idea. Random programming questions via ChatGPT.com maybe as it was pre codex.
I think I gave Plexamp (app) an API token, that’s the only thing I’ve ever come up with on my own speculation
Can you sign up again under a new account? Does it let you use it this way?
That's pretty interesting. It was already well-known that you could easily remove safety training from open-weights models by a bit of finetuning, but apparently you don't even need a finetuning dataset, as long as you have just a few prompts and another LLM to judge responses? Let's see if the abliteration people take a note of this.
>Submitted on 5 Feb 2026