This is the second of two posts resulting from a recent Astra/MATS research project investigating exploration hacking in AI debate. They are designed to be standalone, but we encourage interested readers to read both. The first focuses on our empirical results, this post focuses on a new conceptual framework. Authors...
This is the first of two posts resulting from a recent Astra/MATS research project investigating exploration hacking in AI debate. They are designed to be standalone, but we encourage interested readers to read both. This post focuses on our empirical results, the second focuses on a new conceptual framework. Authors...
Preamble This post is my attempt to try and organise some thinking about AI alignment in a way that will act as a partial-overview to the core ideas and approaches. It is mostly a review of existing ideas arranged with some light opinions thrown in. I do not expect the...
Introduction This piece seeks to explore an interesting way of defining intelligent systems such that we can theoretically quantify their general intelligence. From this, further tools and ideas for comparing these entities could be developed. The definitions are not meant to be philosophical truths, rather they are meant to be...