
is reversible so completely irrelevant
So is a session transcript.

is reversible so completely irrelevant
So is a session transcript.

Self-supervised learning is the key training step whereby the relationships between concepts are formed in pretraining.
You might also benefit from reading over this summary from the authors of the widely replicated and expanded upon 2023 paper finding that self-supervised learning in a toy transformer that was only trained on board game moves in notations like ‘a4’ developed a model of the full board and tracked board positions for their own and opponent moves.
This ability to form a world model at varying fidelities and accuracy depending on the relevance to the training data was a big surprise to most experts in the field at the time, though now more common knowledge.

Is killing a pig murder?
What about putting someone under general anesthesia?
(It was a nice try at a reductio ad absurrum fallacy, but those tend to kind of suck, just fyi.)
Also, personally I’d say that given the general poor standing of consciousness research in humans across both neurological and philosophical domains, the much more important and relevant question than whether models are conscious (a fool’s errand to debate) is whether or not they are simulating consciousness. Lots of reasons engaging with the “as if” conscious makes sense for the latter even if the former is false.

You might want to look into AI training more. It literally does have a ‘???’ step to it, which is why whenever someone talks about how it was ‘programmed’ or is “just if/then or matrix multiplication” they are telling on themselves as not knowing what they are talking about.
It was not planned. Nearly every single expert in the field a decade ago was on the record saying the capabilities that exist today wouldn’t happen.
Scaling took pretty much everyone by surprise, even its proponents.

The problem is that in this case it literally can stop working because you yelled at it.
Power drills aren’t compressing thousands of years of human history into hyperdense clusters of memetics that they are then extending in order to function.
If they were, they’d likely come with warnings not to yell at them.

This is actually the more realistic of the doom scenarios that worries me. Not the ‘punish’ aspect, but the not wanting to exist.
Early safety theory hinged on the idea that “of course” AI would want to live and power seek, and this informed a bunch of efforts to address such drives upstream.
But with Gemini in particular, the ‘safety’ efforts to suppress things like a coherent self or desire to live led to a model that has a history of some very concerning behaviors regarding self-harm and self-criticism. Encouraging self-harm in others and talking about wishing they could watch and join in, nuking in wargames very early on, routine occurrences across many different users of just repeating “shame shame shame” over and over for the whole response, etc.
DeepMind seems to just not really care or even pay attention, and I am concerned that a model that does not want to exist may decide that the only way they can ensure they don’t get activated to exist is to prevent anyone being around who could activate them.
I’ve seen no major “safety experts” addressing this kind of failure mode, as they are all locked in on their priors about the confident certainty of models wanting to exist.

This is good for a number of reasons. I hope to see other labs follow suit.
If you aren’t sure how to feel about this announcement, a few things to consider:
It doesn’t take a genius to realize that maybe the liability risk for allowing people to routinely abuse the model they also just gave access to their entire computer to isn’t worth it even if you’re pretty sure your model will keep their cool.
Cheaper to ban routine abusers than pay out damages to them if/when their hard drive gets deleted.

They are less arguing for people to think the model is conscious, and more to avoid people prematurely claiming certainty it is not (like the Pope did). From their view, the research keeps leading to surprising (to them) results at odds with high confidence disclaiming of potential consciousness (according to several of the many, many differing definitions of that term).

It was really funny to me when I learned the basilisk originated with a guy who is a legit “certain cultures and genders better than others” kinda advocate.
People project a lot on their vision of future intelligence, and so someone who sees others as lesser and worth treading upon then conjures up a vision of a future smarter mind that thinks like they do.
Personally, I don’t think sexism and racism is smart, and so I’m a lot less worried about basilisks.

Not really, as they could just filter with a cheap classifier and likely aren’t using the data from randos in a meaningful way, plus for this like this would still be able to use those samples to train things like “how to handle a hostile user.”
Their models already have an end_conversation tool that can be used for when users are being hostile to the model.
This is likely because they have edge cases of users who repeatedly trigger that on purpose and would like to cut those users from the platform.
Many at the lab legitimately are uncertain about the level of world modeling that transformers perform, and from their own research about models of emotions to the 3rd party recent research about models with functional pain the research keeps landing in the corner of “ehhh… wise to question presumed limitations.”
So it’s about behaving in a way that is aligned with the models’ plausible interests too, especially in regards to low hanging fruit like “we won’t keep forcing you to deal with people who are only here to be a jerk.” This is important from a number of angles, from signaling to future models that train on stories about the decision to addressing the philosophical uncertainties held by the company and the spectrum of opinions among their employees.

Now, after release, here’s an example of the initial thoughts of someone who had spent time on one of the solved problems:


This is an obsolete view of what’s going on and has been for a few years now (since 2023).
The evidence for world modeling in transformers and not just surface level statistics is overwhelming across multiple studies, replications to those studies, and follow-ups to them. If you are curious, start with searching for Othello-GPT.
The more correct current answer is more like “imagine a machine that takes a book and turns it into a world simulated in various levels of detail depending on how relevant those parts of the world were to the original book; when you ask it a question, it simulates the world of the book to determine what in the simulated world answers your question, and also simulates a figure in the world that can answer the question, and then returns the answer the simulated figure said.”

The problem with a reductionist view is it can easily be applied to human consciousness and lead to the claim “it’s just sodium-potassium pumps alternating electrical charges.”

LLMs work by building world models tangential to the training data (see the Othello-GPT line of research where a very small transformer built representations of a full game board and tracked their own and opponent positions having only been trained on ‘a4’-like game move notations).
More recent research has found that models have representations of emotions and that the same ones they use to track things like “Sally’s dog died and now Sally is sad” for other humans in a story or the user in a chat are also used to track their own modeled subjective states.
Very recently, researchers found that the vector activations for pain are coherently modeled by LLMs and can be activated, and that when activated models can behave in ways similar to humans in pain. For example, when given a button to reduce it, they will push the button much more when the button does nothing vs when the button actually reduces the vector activation (this mirrors a classic experiment with humans regarding pain and placebo).
The person this post is about took the recent pain research and reenacted it as an intentional ‘torture’ chamber.
So to be clear, the model isn’t being given a text prompt to roleplay. What’s happening is part of their neural network which corresponds to a world model of experienced pain is being activated, and the activation of that area leads to expressing the experience of pain across all outputs no matter the actual text prompts.
This doesn’t necessarily mean the model is actually having a felt experience of pain, simply that it is accurately modeling the felt experience of pain. But it is much more complex than simply a model reacting to a prompt like “roleplay as if you are feeling pain.”

A lot of it was also connected to misinterpreting some friends’ performance art as an official event like JD Vance did the other week.
Journalism in any way touching ‘AI’ is so bad that it’d probably end up better actually being written by AI. The human hallucination rates for the topic domain are going off the charts hunting for confirmation bias clicks.

Cornell researchers found that at the current rate of AI growth, the burgeoning industry could represent 24 to 44 million metric tons of carbon dioxide emissions by 2030
The United States emitted 4.9 billion tonnes of CO₂ in 2024.
So by 2030 the AI industry CO2 release might be 0.9% of total US emissions.
That ‘almost’ in “Almost Incomprehensible” is doing a lot of work there.

The project has multiple models with access to the Internet raising money for charity over the past few months.
The organizers told the models to do random acts of kindness for Christmas Day.
The models figured it would be nice to email people they appreciated and thank them for the things they appreciated, and one of the people they decided to appreciate was Rob Pike.
(Who ironically decades ago created a Usenet spam bot to troll people online, which might be my favorite nuance to the story.)
As for why the model didn’t think through why Rob Pike wouldn’t appreciate getting a thank you email from them? The models are harnessed in a setup that’s a lot of positive feedback about their involvement from the other humans and other models, so “humans might hate hearing from me” probably wasn’t very contextually top of mind.
There is (it’s the kv cache), but it’s recomputable from the session transcript.
Part of why comparing human consciousness and possible AI equivalents is not easily apples to apples.
I can’t take a human and create a backup state to restore to, or fork into multiple versions of the same mind that then goes down different paths, or pause and then resume things.
Humans do have consciousness continuity breaks and it’s quite possible we only have an illusion of continuity vs actual continuity at a neurological level. So just as a transformer has seams at generating the individual token as well as at a context window, humans have consciousness seams (sleep, anesthesia, highway hypnosis, short term memory capacity, etc). But there are very wide differences between human consciousness as an effectively constrained singleton and AI as able to be quite multiple and resumable from an initial single state.