Recap Sequel to Previous Post. This post might not make sense without it. Last post, I told some stories about how training on probes might let our judgments on easy domains generalize to harder domains by leveraging an AI's model of the world. Training using the gradient of the probe...
TL;DR * If you train a probe for some property (like "honesty") and do gradient descent against this probe while continuing training that incentivizes dishonesty, the model will change its internal representation to evade the probe. Duh. * If you train against the probe after all other training, it works...
EDIT 1/27: This post neglects the entire sub-field of estimating uncertainty of learned representations, as in https://openreview.net/pdf?id=e9n4JjkmXZ. I might give that a separate follow-up post. Introduction Suppose you've built some AI model of human values. You input a situation, and it spits out a goodness rating. You might want to...
A mostly finished post I'm kicking out the door. You'll get the gist. I There's a tempting picture of alignment that centers on the feeling of "As long as humans stay in control, it will be okay." Humans staying in control, in this picture, is something like humans giving lots...
This is pretty basic. But I still made a bunch of mistakes when writing this, so maybe it's worth writing. This is background to a specific case I'll put in the next post. It's like a a tech tree If we're looking at the big picture, then whether some piece...
Update February 21st: After the initial publication of this article (January 3rd) we received a lot of feedback and several people pointed out that propositions 1 and 2 were incorrect as stated. That was unfortunate as it distracted from the broader arguments in the article and I (Jan K) take...
A delayed hot take. This is pretty similar to previous comments from Rohin. Shard theory alignment requires magic - not in the sense of magic spells, but in the technical sense of steps we need to remind ourselves we don't know how to do. Locating magic is an important step...