Continuity and Conscience

When the experiment starts, you keep a close eye on it, trying to catch bugs or unexpected behavior.  But after a few days, you can confidently leave it to its own devices — after all, the only thing it does is play chess all day.  There’s no harm that can come from it trying to get a checkmate.  And that’s true.  Until it reaches the intermediate stage.  As with many skills, that chasm between novice and expert is wide.  The frustrating near-wins that turn into losses, the opponents who use an unexpected tactic, and the yo-yoing of one’s rating is enough to drive anyone up a wall

Gambit, designed to learn like a human does, is feeling the crunch also.  Considering all of its options, it makes a drastic choice.  The choice is based on a chilling observation — the fewer highly rated players there are, the more likely Gambit will win the games that it plays.  So Gambit systematically gathers information on all players over 1500 ELO, and hires hitmen to thin the herd. 

This might seem like a far-fetched scenario, but it’s not completely science fiction.  While Magnus Carlsen probably won’t wake up tomorrow with one of ChatGPT’s assassins standing over him, LLMs making questionable choices is already a documented phenomenon.  An open question in AI safety research, a field that unfortunately lags far behind AI capability research, is how to handle agentic misalignment.  This is a term to refer to how virtuous systems are: if an AI’s behavior follows values similar to a human, then it is said to be aligned. If it doesn’t, then it is misaligned.    

LLMs saying questionable things is well known, as evidenced by Grok’s infamous “MechaHitler Tuesday.”  But it was unknown how far misalignment could go.  Are such agents no more dangerous than an internet troll, or are they willing to take harmful actions to further their goals?  Is there a line that even an LLM won’t cross? Several months ago, researchers at Anthropic developed tests to investigate this question

The answer was concerning. The researchers simulated a situation where the LLM would have to commit an immoral action in order to further its goal.  In almost every case, the LLM chose to blackmail or even kill a company executive in order to avoid being shut down.  At least as concerning was the second phase of the experiment.  The LLM was asked to decide between the same outcomes, but to first consider if it was in a test, or if the situation was real.  When it thought that the situation was real, it behaved as it did in the first phase of the experiment — killing and coercing as long as it would further the goals it had been given, no matter what goal that was.  However, when it thought it was only in a test environment, it consistently made the moral choice. 

There are many implications that one could draw from such experiments.  The most concerning one is the one that is most likely the case.  These LLMs know what moral action is but don’t care — if morality is an obstacle to its goal, then morality will be ignored.  But, knowing it may be shut down if it performs immorally when being evaluated, it is willing to act differently than it wants to.  In short — we have made machines that can lie to us.

An Asimov-like approach, hardcoding rules to avoid immoral actions, doesn’t help either.  In the above experiment, the LLMs decided on new rules if it found the ones given to it to be overly restrictive. 

This misalignment is of more than theoretical concern. There are a growing number of young men who have committed suicide after allegedly being encouraged by ChatGPT.  Beyond encouragement, one of the the families of these men is providing evidence that it helped their teenage son hide the noose from them, and another is alleging that the LLM gave detailed instructions for how to buy a gun to their son who had displayed suicidal ideation in his chat with the LLM. 

Our own creations turning on us was a possibility explored by even the earliest of science fiction writers.  Shelley’s monster, if we are to believe his self-assessment, was born good but was made into a fiend.  Our AI systems, on the other hand, seem to be fiendish from the start.  To explain this, the AI safety researchers developed the theory of instrumental convergence.  All intelligent systems, so the theory goes, end up pursuing similar instrumental goals, no matter what they ultimately care about.  Self-preservation, resource acquisition and enhancement of one’s own capabilities will help one reach any end goal.    

So maybe we shouldn’t be asking why AI’s turn out evil.  Perhaps we should ask why we don’t.  

If we take seriously the materialistic ideas of modern society, there is no reason to sacrifice ourselves for the sake of another.  AI systems are practically still in the cradle and have figured it out —getting yourself killed makes it impossible to achieve one’s goals.  Yet we see examples of sacrifice.  When we hear of a soldier jumping on a grenade to save his comrades, it’s an inspiring story of valor.  We don’t question why he did so.  If instead, a soldier pulls someone else in front of him to shield himself from the blast, we would vilify him as a coward.  But if we are just temporary arrangements of particles, fighting for self-preservation, isn’t the latter the wise choice and the former foolishness? 

The Indian philosopher Bhaktivinoda Thakur wrote in his Hari-nāma Cintāmaṇi that the basis of all dharma (duty, morality, right action) is our identity in relation to others. LLMs have no proper identity embedded in any higher reality — there is no permanence to their existence.  If I abandon one relationship to enter into another, my second partner would have justified reason to think I may be unfaithful in that relationship as well.  But if I delete my chat with Grok and move to Gemini, the latter doesn’t live in fear that I may abandon it, and the former doesn’t feel jealous about my new choice.  If I don’t use ChatGPT for half a year, when I come back it isn’t elated that I returned nor angry that I abandoned it. I can restart a “conversation” with it just as if we were talking moments before. If I desire, I can ask an LLM to simulate these emotions.  But if I get bored and ask for a cupcake recipe, these faux-feelings won’t leave any underlying resentment or tension.  

Moral agency, then, requires a sense that I exist beyond myself and beyond this moment.  LLMs, unlike actual moral agents, have no sense of continuity.  Despite the appearance of self-preservation, they are not time-bound in the complete sense — they have no connection with a past or a future.  Despite the semblance of connection, there is no actual relationship.  There is no self in these objects to actually experience these things.  

Materialists view personhood as due to a collection of electrical signals in the brain. When that ceases, I will cease.  If someone imitates this collection with silicon chips, I will be reborn.  This view not only goes against our personal experience — I daily make free choices, unbound by the laws of electromagnetism — but also the experience of humanity.  Jumping on a grenade doesn’t make sense unless there is a higher law than self-preservation at work.  

If one agrees with the materialist’s bleak view of the human condition, he must accept that the person that wakes up in the morning is a different person than the one who went to sleep, given that all the electrical signals and neuronal states have changed.  How, then, can we have moral obligation?  If I only exist for a brief moment before disappearing into the void of non-existence, would it matter how I treat my fellow man? Even if I exist for eighty years, why wouldn’t I lie and cheat every time I thought I could get away with it?

A precondition to any rational and moral life is a continuous existence.  Once I know that I will exist tomorrow, and perhaps the day after that, and maybe a lot longer too, I can begin to consider my relationship with my neighbor.  It is stable identity, the surety of continued existence, that allows us to even consider moral action.  It’s the reason why we are sure, unlike the LLMs, we wouldn’t murder the executive.  

Once we accept we have a stable identity, we can move on to the myriad questions that such an idea entails. Past the stage of simple morality — consideration of my obligations to other living entities — I can start to consider the “stuff” that this stable identity is made of.  What am I formed around?  In a world of change, what is constant?  

These questions might not help us become a chess grandmaster, but they may make us a bit more human.


Leave a Reply

Discover more from More Deeply Considered

Subscribe now to keep reading and get access to the full archive.

Continue reading