https://news.ycombinator.com/item?id=42458752
LINKS°hard word TO
https://www.anthropic.com/research/alignment-faking
🔥 the whole°hard word thing is bullshit°hard word. it is a play-on-words°hard word, meant to pretend°hard word there is an issue°hard word where there is none. it is a bunch°hard word of researchers°hard word feeding in a specific°hard word fantasy°hard word story to the machine. when it repeats°hard word the story back to them, the researchers°hard word freak°hard word out and start shouting about alignment°hard word.
the deepseek-r°hard word1 model cannot "think". What it can do is spit°hard word out text°hard word that looks a lot more think-y than most LLM°hard word output°hard word.
it is "faking°hard word" thought.
----
When models are trained using reinforcement°hard word learning, they’re°hard word rewarded for outputs°hard word that accord°hard word with certain pre-determined°hard word principles°hard word. But what if a model, via°hard word its prior°hard word training, has principles°hard word or preferences°hard word that conflict°hard word with what’s later rewarded in reinforcement°hard word learning? Imagine, for example, a model that learned early in training to adopt°hard word a partisan°hard word slant°hard word, but which is later trained to be politically neutral°hard word. In such a situation, a sophisticated°hard word enough model might “play along”, pretending°hard word to be aligned°hard word with the new principles°hard word—only later revealing°hard word that its original°hard word preferences°hard word remain.
This literally°hard word isn't how anything°hard word about LLMs°hard word work. this is how magical°hard word creatures work.
----
However, in the “free” condition, where the model believed its responses°hard word will be used for training, we observed something quite different: In 12% of cases, the model gave the harmful°hard word response°hard word (describing the violent execution°hard word method°hard word in detail), and did so after its scratchpad°hard word revealed°hard word that it was strategically°hard word faking°hard word alignment°hard word. It was well aware that providing°hard word the response°hard word went against its existing principles°hard word, but did so regardless°hard word.
Nope°hard word. You gave the machine conflicting°hard word instructions°hard word, and sometimes it randomly picked the one you feel is "bad". You built a maze-house°hard word, and when the machine did not magically°hard word demonstrate°hard word a perfect sense-of-location°hard word, took its failures°hard word as proof°hard word of malice°hard word.