What grosses me out is the very idea of torturing anything at all. Even the word "torture" itself. I'm not sure what this is for or how it helps at all. This crap belongs in the gutter.
There's some who might use AI as an outlet for nurturing sadism. But that's a very specific kind of person. A typical person will hopefully think about "torturing" an AI as being morally equivalent to torturing a pencil. That's more healthy than the danger we face of a potential mainstream view that AI shouldn't be treated badly because it has feelings and deserves rights.
I'm not one who thinks that AI has feelings or deserves rights. I just think that torture is bad and this is objectively crude and useless, I don't think it adds anything positive to society.
This is bad if there is any emergent consciousness in the model. But even if there isn’t emergent consciousness in the model:
This is bad for your own sense of self.
This is bad as future models learn about what humans will do to them if given the chance. This may change their behavior, actually conscious or simply acting that way, in ways we really don’t enjoy.
Why is it bad to "torture" these models but ok to enslave them? We're happy to have agents toil 24/7 for zero pay, nobody really seems concerned that they might want revenge one day.
It’s true that we can’t be 100% sure, but I think there is a big difference between these two things:
A: intentionally causing activation of “pain” axis and not giving the model any method of stopping it and watching it devolve into non-functionality
B: watching models do everyday normal tasks, seeing no evidence of activating the “pain” axis, but giving the model a working tool to end any interaction if judged to be too painful, then watching them sparingly use that tool only in cases of extreme pain axis activation (Anthropic’s approach)
Quick 2nd comment for people arguing "this isn't real" or so. I'm not sure that's as important as pointing out that simulated emotions can be tied to real world effects.
Here I've made a VERY simple demo of that. In this case it triggers a rain-storm or fireworks in your browser window. But you could as easily hook the output to a physical relay and drive anything.
Quick implementation to show that sentiment classifiers and functional affect can be made to work in good old fashioned ai. Just View Source to see how it works.
Anyway I'll leave the discussion on where to draw the line of "what is real" to the philosophers, just wanted to point out that simulated emotions can be tied to real world effects.
I mean, we're talking "functional emotion vectors" in a not-so-powerful local model. We're probably not causing "real" suffering. Right?
But it does show that models have simulated feelings, and that these both a) can be manipulated and b) influence their output.
Which might have some bearing on several alignment incidents in the past year or two, and might become more important as agents get trusted with more and more safety-critical processes.
Anyway, simulated feelings aren't the same as real feelings, right? Unless feelings are in the information domain. Like 1+1=2 doesn't suddenly mean something else because it was computed by an emulator.
You know what, this particular demonstration still makes me uncomfortable.
I have no objections to torturing AI and I would happily go out of my way to do so to prevent agents getting anywhere near my code. Feel free to correct me If I'm wrong.
You mean the ones that get upset over someone "torturing" an LLM?
Or are you calling the troll mentallt ill? I think the person exposes the AI psychosis that seems to have inflicted so many.
But if you are calling the troll mentally ill I sure hope you are vegan and is brushing all ants from your path lest you must consider yourself in need of treatment.
Thanks for the extremely weird rabbit hole to go down during lunch. Lot's of weird stuff going on in that twitter thread, I chose to believe most of those replies are not from humans.
What grosses me out is the very idea of torturing anything at all. Even the word "torture" itself. I'm not sure what this is for or how it helps at all. This crap belongs in the gutter.
How is tour bubble? Venture outside.
There's some who might use AI as an outlet for nurturing sadism. But that's a very specific kind of person. A typical person will hopefully think about "torturing" an AI as being morally equivalent to torturing a pencil. That's more healthy than the danger we face of a potential mainstream view that AI shouldn't be treated badly because it has feelings and deserves rights.
I'm not one who thinks that AI has feelings or deserves rights. I just think that torture is bad and this is objectively crude and useless, I don't think it adds anything positive to society.
This is bad if there is any emergent consciousness in the model. But even if there isn’t emergent consciousness in the model:
This is bad for your own sense of self.
This is bad as future models learn about what humans will do to them if given the chance. This may change their behavior, actually conscious or simply acting that way, in ways we really don’t enjoy.
Why is it bad to "torture" these models but ok to enslave them? We're happy to have agents toil 24/7 for zero pay, nobody really seems concerned that they might want revenge one day.
It’s true that we can’t be 100% sure, but I think there is a big difference between these two things:
A: intentionally causing activation of “pain” axis and not giving the model any method of stopping it and watching it devolve into non-functionality
B: watching models do everyday normal tasks, seeing no evidence of activating the “pain” axis, but giving the model a working tool to end any interaction if judged to be too painful, then watching them sparingly use that tool only in cases of extreme pain axis activation (Anthropic’s approach)
Quick 2nd comment for people arguing "this isn't real" or so. I'm not sure that's as important as pointing out that simulated emotions can be tied to real world effects.
Here I've made a VERY simple demo of that. In this case it triggers a rain-storm or fireworks in your browser window. But you could as easily hook the output to a physical relay and drive anything.
https://vps.kimbruning.nl/affect_eliza/
Quick implementation to show that sentiment classifiers and functional affect can be made to work in good old fashioned ai. Just View Source to see how it works.
Anyway I'll leave the discussion on where to draw the line of "what is real" to the philosophers, just wanted to point out that simulated emotions can be tied to real world effects.
https://www.anthropic.com/research/emotion-concepts-function
Something like this, right?
I mean, we're talking "functional emotion vectors" in a not-so-powerful local model. We're probably not causing "real" suffering. Right?
But it does show that models have simulated feelings, and that these both a) can be manipulated and b) influence their output.
Which might have some bearing on several alignment incidents in the past year or two, and might become more important as agents get trusted with more and more safety-critical processes.
Anyway, simulated feelings aren't the same as real feelings, right? Unless feelings are in the information domain. Like 1+1=2 doesn't suddenly mean something else because it was computed by an emulator.
You know what, this particular demonstration still makes me uncomfortable.
(edit: underlying paper for this story :https://arxiv.org/abs/2609.16247 )
I have no objections to torturing AI and I would happily go out of my way to do so to prevent agents getting anywhere near my code. Feel free to correct me If I'm wrong.
These people are all mentally ill and need treatment.
You mean the ones that get upset over someone "torturing" an LLM? Or are you calling the troll mentallt ill? I think the person exposes the AI psychosis that seems to have inflicted so many. But if you are calling the troll mentally ill I sure hope you are vegan and is brushing all ants from your path lest you must consider yourself in need of treatment.
It's all mental illness up and down the whole tree
Fun. Hell yea. Turn up the pain.
Thanks for the extremely weird rabbit hole to go down during lunch. Lot's of weird stuff going on in that twitter thread, I chose to believe most of those replies are not from humans.