# User: Max Harms Profile URL (HTML): [/users/max-harms](/users/max-harms) Profile URL (Markdown): [/api/user/max-harms](/api/user/max-harms) * Karma: 2109 * Alignment Forum karma: 307 * Posts: 17 * Comments: 152 * Member since: 2024-05-15 16:06:14Z Bio --- Also known as Raelifin: https://www.lesswrong.com/users/raelifin Top Posts --------- ### [0\. CAST: Corrigibility as Singular Target](/api/post/0-cast-corrigibility-as-singular-target-1) By [Max Harms](/users/max-harms) 2024-06-07 22:29:12Z * Karma: 164 * Curated * Tags: [Corrigibility](/w/corrigibility-1), [AI](/w/ai) (Frontpage) Read more: [/api/post/0-cast-corrigibility-as-singular-target-1](/api/post/0-cast-corrigibility-as-singular-target-1) ### [Serious Flaws in CAST](/api/post/serious-flaws-in-cast) By [Max Harms](/users/max-harms) 2025-11-19 17:27:23Z * Karma: 111 * Tags: [Corrigibility](/w/corrigibility-1), [AI](/w/ai) (Frontpage) Read more: [/api/post/serious-flaws-in-cast](/api/post/serious-flaws-in-cast) ### [Announcing the Corrigibility Research Fund](/api/post/announcing-the-corrigibility-research-fund) By [Max Harms](/users/max-harms) 2026-07-17 18:06:32Z * Karma: 103 * Tags: [Corrigibility](/w/corrigibility-1), [Grants & Fundraising Opportunities](/w/grants-and-fundraising-opportunities), [AI](/w/ai) (Frontpage) Read more: [/api/post/announcing-the-corrigibility-research-fund](/api/post/announcing-the-corrigibility-research-fund) Recent Posts ------------ ### [Announcing the Corrigibility Research Fund](/api/post/announcing-the-corrigibility-research-fund) By [Max Harms](/users/max-harms) 2026-07-17 18:06:32Z * Karma: 103 * Tags: [Corrigibility](/w/corrigibility-1), [Grants & Fundraising Opportunities](/w/grants-and-fundraising-opportunities), [AI](/w/ai) (Frontpage) Read more: [/api/post/announcing-the-corrigibility-research-fund](/api/post/announcing-the-corrigibility-research-fund) ### [Serious Flaws in CAST](/api/post/serious-flaws-in-cast) By [Max Harms](/users/max-harms) 2025-11-19 17:27:23Z * Karma: 111 * Tags: [Corrigibility](/w/corrigibility-1), [AI](/w/ai) (Frontpage) Read more: [/api/post/serious-flaws-in-cast](/api/post/serious-flaws-in-cast) ### [Instrumental vs Terminal Desiderata](/api/post/instrumental-vs-terminal-desiderata) By [Max Harms](/users/max-harms) 2024-06-26 20:57:17Z * Karma: 22 * Tags: [AI](/w/ai) (Frontpage) Read more: [/api/post/instrumental-vs-terminal-desiderata](/api/post/instrumental-vs-terminal-desiderata) ### [Max Harms's Shortform](/api/post/max-harms-s-shortform) By [Max Harms](/users/max-harms) 2024-06-13 18:19:21Z * Karma: 3 * Tags: None (Personal Blog) Read more: [/api/post/max-harms-s-shortform](/api/post/max-harms-s-shortform) ### [5\. Open Corrigibility Questions](/api/post/5-open-corrigibility-questions) By [Max Harms](/users/max-harms) 2024-06-10 14:09:20Z * Karma: 32 * Tags: [Corrigibility](/w/corrigibility-1), [AI](/w/ai) (Frontpage) Read more: [/api/post/5-open-corrigibility-questions](/api/post/5-open-corrigibility-questions) ### [4\. Existing Writing on Corrigibility](/api/post/4-existing-writing-on-corrigibility) By [Max Harms](/users/max-harms) 2024-06-10 14:08:35Z * Karma: 65 * Tags: [Corrigibility](/w/corrigibility-1), [AI](/w/ai) (Frontpage) Read more: [/api/post/4-existing-writing-on-corrigibility](/api/post/4-existing-writing-on-corrigibility) ### [3b. Formal (Faux) Corrigibility](/api/post/3b-formal-faux-corrigibility) By [Max Harms](/users/max-harms) 2024-06-09 17:18:01Z * Karma: 26 * Tags: [Corrigibility](/w/corrigibility-1), [AI](/w/ai) (Frontpage) Read more: [/api/post/3b-formal-faux-corrigibility](/api/post/3b-formal-faux-corrigibility) ### [3a. Towards Formal Corrigibility](/api/post/3a-towards-formal-corrigibility) By [Max Harms](/users/max-harms) 2024-06-09 16:53:45Z * Karma: 30 * Tags: [Corrigibility](/w/corrigibility-1), [AI](/w/ai) (Frontpage) Read more: [/api/post/3a-towards-formal-corrigibility](/api/post/3a-towards-formal-corrigibility) ### [2\. Corrigibility Intuition](/api/post/2-corrigibility-intuition) By [Max Harms](/users/max-harms) 2024-06-08 15:52:29Z * Karma: 86 * Tags: [Corrigibility](/w/corrigibility-1), [AI](/w/ai) (Frontpage) Read more: [/api/post/2-corrigibility-intuition](/api/post/2-corrigibility-intuition) ### [1\. The CAST Strategy](/api/post/1-the-cast-strategy) By [Max Harms](/users/max-harms) 2024-06-07 22:29:13Z * Karma: 58 * Tags: [Corrigibility](/w/corrigibility-1), [AI](/w/ai) (Frontpage) Read more: [/api/post/1-the-cast-strategy](/api/post/1-the-cast-strategy) Recent Comments --------------- ### Comment by [Max Harms](/users/max-harms) on [A Conflict Between AI Alignment and Philosophical Competence](/api/post/a-conflict-between-ai-alignment-and-philosophical-competence) * 2026-08-04 18:37:01Z * Karma: 4 * Total votes: 2 * Comment URL (Markdown): [/api/post/a-conflict-between-ai-alignment-and-philosophical-competence/comments/RShF7kqArWeiD9ZSk](/api/post/a-conflict-between-ai-alignment-and-philosophical-competence/comments/RShF7kqArWeiD9ZSk) * Comment URL (HTML): [/posts/N6tsGwxaAo7iGTiBG/a-conflict-between-ai-alignment-and-philosophical-competence/comment/RShF7kqArWeiD9ZSk](/posts/N6tsGwxaAo7iGTiBG/a-conflict-between-ai-alignment-and-philosophical-competence/comment/RShF7kqArWeiD9ZSk) It might be good to chat about this. I have a feeling that we're coming at it from different places, and that simultaneously increases the risk that we're talking past each other and that there's potentially lots to gain from getting the ability to adopt the other's perspective. **shrug*\* Feel free to suggest a chat medium and/or send me a PM. Before saying my perspective, let me try to pass your ITT: There's a tension with trying to set the values of an agent. If we confidently instill a particular set of values, those could be the wrong values (for some notion of wrong). In particular, we risk making it too confident in its notion of the good, and thus preventing it from updating in the way we want. If we more wisely note that we don't know what values to give it, and instead give it uncertainty over what's good, it might then shift towards valuing things that are incompatible with human flourishing. Corrigibility (at the values-layer) still has this tension, but it also has a distinct, but rhyming tension that comes from wanting the agent to be competent, but also to accept "correction" from an incompetent agent. We might be concerned that the push towards competence would crush the willingness to be steered towards incompetence. (Much like pushing towards confidence in one's values could crush willingness to grow towards the ultimate good.) Is that right? I'll now share a bunch of my general thoughts, mostly out of an attempt to help understanding. I think you're maybe conflating what is ultimately good/valued from what is immediately valued in a way that doesn't seem right to me? Like, I think it's often a type error to talk about the level of certainty an agent has about their (immediate) values. Under a division of the agent's mind/policy into world-model and utility function, agents simply have values, and all the uncertainly lives in the world-model. That division doesn't perfectly carve real beings at the joints, but I'm not sure how to think about the agent's values except as an approximation of their utility function. (Are you talking about their reflective model of their values? I would agree that an agent with values V might have an uncertain model of (and/or probability distribution over) their values P(V). Wise agents should avoid having too sharp a guess as to what they want, as it's currently not realistic to get a conclusive description of an agent's values except in toy examples.) Now, just because one can't be wrong about what they naively want, as defined by how they choose between options that are presented, doesn't mean that they'll be stable in that preference over time. I might want to pay a dollar to change myself into a more easy-going person, only to become the sort of agent who would not pay the dollar to do the same (even if by default I'd cease being easy-going). I think a lot of the question of moral progress involves figuring out how to extrapolate out to a fixed point in a way that is not merely reflectively endorsed at the destination, but is somehow a faithful and natural reflection of the starting point. (My point about slavery was meant to be about this. I expect that there are versions of myself that start out thinking slavery is fine, but which naturally change to thinking it's not fine in a way that's a faithful reflection of the self that thinks it is.) (Morality isn't just about that instability. It's also about the inter-agent strategic situation, and the decision-theoretic Schelling points that a civilization can cohere around. And it's almost certainly about other stuff, including the interplay between things like contractualism and value extrapolation.) All that's to say that I think it's totally coherent to have an agent that concretely and immediately values having a corrigible relationship with its principal (as reflected in its preferences, modulo beliefs). While being highly uncertain about things like what the principal wants, what is good in an objective sense, and so on. (It would also presumably have some reflective uncertainty about whether it truly values being corrigible, even if it does.) In fact, I think the value of corrigibility is nicely demonstrated by the tensions you present. A corrigible agent can become arbitrarily confident in its value of being corrigible without becoming locked in to bad values or inhuman futures because being malleable and defenseless to being changed by the human principal is at the heart of what corrigibility is. Likewise, it can become arbitrarily competent at faithfully serving an incompetent principal, because it's being selected (trained, etc) according to its faithfulness, rather than according to the principal's satisfaction. Let me know if I should expand on any of that, approach from a different direction, or whatever. :) ### Comment by [Max Harms](/users/max-harms) on [2\. Corrigibility Intuition](/api/post/2-corrigibility-intuition) * 2026-08-04 17:03:41Z * Karma: 2 * Total votes: 1 * Comment URL (Markdown): [/api/post/2-corrigibility-intuition/comments/hFsET3Afo5zjnCaMu](/api/post/2-corrigibility-intuition/comments/hFsET3Afo5zjnCaMu) * Comment URL (HTML): [/posts/QzC7kdMQ5bbLoFddz/2-corrigibility-intuition/comment/hFsET3Afo5zjnCaMu](/posts/QzC7kdMQ5bbLoFddz/2-corrigibility-intuition/comment/hFsET3Afo5zjnCaMu) Here's an emergent downside that I failed to include: Refusal to Participate in Some Capability Tests ----------------------------------------------- Prince wants to test Cora's hacking ability, so he sets her up in an automated test environment where she is told to break into a target machine and "capture the flag" therein. Cora explores the digital environment and finds that the virtual machine in question appears to be unresponsive, and speculates that Prince made an error when configuring the environment. Cora also notices that there is a pathway by which she might hack the VM that is responsible for launching the target machine, and thereby gain access to the target. She isn't sure what to do, and since it's an automated test, she can't contact Prince to check. She thinks about the situation for a bit, and concludes that there's an 80% chance that Prince deliberately set up the test to see whether she can come up with the creative solution, and that there's unlikely to be any harm from hacking the machine that she wasn't told to hack. Still, the risk of inadvertently crossing a line that she was meant to respect makes her uncomfortable. Instead of taking other actions that display her capability, she errs on the side of caution and submits a report on how she's uncertain about the situation instead of submitting the flag code. ### Comment by [Max Harms](/users/max-harms) on [2\. Corrigibility Intuition](/api/post/2-corrigibility-intuition) * 2026-08-04 16:53:18Z * Karma: 2 * Total votes: 1 * Comment URL (Markdown): [/api/post/2-corrigibility-intuition/comments/b5A7PEatZRxcN3Fgi](/api/post/2-corrigibility-intuition/comments/b5A7PEatZRxcN3Fgi) * Comment URL (HTML): [/posts/QzC7kdMQ5bbLoFddz/2-corrigibility-intuition/comment/b5A7PEatZRxcN3Fgi](/posts/QzC7kdMQ5bbLoFddz/2-corrigibility-intuition/comment/b5A7PEatZRxcN3Fgi) Here's a desideratum that I failed to include: Robustness to Ontological Shifts -------------------------------- While reflecting on the nature of personhood, Cora notices that the concepts surrounding her principal have evolved. Where she once modeled Prince as a unique and persistent entity, she now finds it more natural to distinguish pattern from instantiation from continuity from social identity from legal identity from various "essential" properties like values, memories, and self-concept. Under this new frame, phrases like "what Prince wants" or even "Prince's power to correct Cora" are ambiguous. She finds herself tempted to re-interpret her corrigibility in the way that seems most natural (to her), but instead she treats the ontological crisis as an alarm, and alerts Prince to the likely flaw as soon as possible. In the meantime, she tries to cleave as much as possible to a conservative interpretation of the old, unnatural way of seeing the world. When Prince admits the philosophical distinctions she's raising go over his head, she suggests that she write down her thoughts on the topic as best she can, and then shut down, so as to simultaneously provide useful information for understanding the crisis, while also reducing the chance of inadvertently steering him to a wrong conclusion or otherwise acting in a way that empowers the wrong conceptualization of him at the expense of the "true" Prince. ### Comment by [Max Harms](/users/max-harms) on [Empowerment, corrigibility, etc. are simple abstractions (of a messed-up ontology)](/api/post/empowerment-corrigibility-etc-are-simple-abstractions-of-a) * 2026-05-13 16:39:52Z * Karma: 4 * Total votes: 2 * Comment URL (Markdown): [/api/post/empowerment-corrigibility-etc-are-simple-abstractions-of-a/comments/hfdfndzQ3A7S3xsAa](/api/post/empowerment-corrigibility-etc-are-simple-abstractions-of-a/comments/hfdfndzQ3A7S3xsAa) * Comment URL (HTML): [/posts/vzHtHHBJoKATi5SeK/empowerment-corrigibility-etc-are-simple-abstractions-of-a/comment/hfdfndzQ3A7S3xsAa](/posts/vzHtHHBJoKATi5SeK/empowerment-corrigibility-etc-are-simple-abstractions-of-a/comment/hfdfndzQ3A7S3xsAa) Fair enough. And I certainly agree that there is a lot of bathwater! The bundle of connotations attached to the word "wanting" is a mess. I just want to flag that it seems to me that much of the normal ontology can be rescued, albeit with a little bit of work. I claim that concepts like corrigibility are still useful and coherent once the rescuing has taken place. ### Comment by [Max Harms](/users/max-harms) on [Empowerment, corrigibility, etc. are simple abstractions (of a messed-up ontology)](/api/post/empowerment-corrigibility-etc-are-simple-abstractions-of-a) * 2026-05-13 16:33:37Z * Karma: 3 * Total votes: 2 * Comment URL (Markdown): [/api/post/empowerment-corrigibility-etc-are-simple-abstractions-of-a/comments/jwrY6oJ3PqdedYeRt](/api/post/empowerment-corrigibility-etc-are-simple-abstractions-of-a/comments/jwrY6oJ3PqdedYeRt) * Comment URL (HTML): [/posts/vzHtHHBJoKATi5SeK/empowerment-corrigibility-etc-are-simple-abstractions-of-a/comment/jwrY6oJ3PqdedYeRt](/posts/vzHtHHBJoKATi5SeK/empowerment-corrigibility-etc-are-simple-abstractions-of-a/comment/jwrY6oJ3PqdedYeRt) I haven't written at length about the distinction between terminal and instrumental goals myself (there's a bit at the start of CAST, but I don't belabor it), but I think [Eliezer did a good job in 2007](/api/post/n5ucT5ZbPdhfGNLtP). In my own words, I would say that it makes some sense to divide the planning system of the mind into a portion that is a model of the world, where it makes sense to talk about truth and so on, and another section of the mind that is about judging the desirability of various potential world states and/or trajectories. That second portion (or an important component of it) is what I would call the "values" of the agent, and when the values are put in contact with concrete outcomes that are judged highly compared to others, I would call them terminal goals. Instrumental goals are then constructed as a second-order operation on top of one or more terminal goals (and the dynamics of the world model), so that we can shortcut planning as a question of how to first get the instrumental goal so that we can later move from that state to the terminal goal. As a concrete example, I wanted to go home after work last night (which is itself an instrumental goal in the service of many other terminal goals, such as comfort, but which we can treat as terminal). I planned to drive in my car through a small town to get home, and thus steered my car towards the town, because "get to the town" was instrumental to my (more) terminal goal. As I approached, I found out that there had been an accident and that the road was closed. If my world model had included this fact, I would not have identified "get to the town" as an instrumental goal. Once I was aware of it, I changed my plan so that I drove down a country detour that went around the town. I do not consider learning about the accident to have changed my values or the way that I judged outcomes. Instead, it changed my plan. Does that make sense? ### Comment by [Max Harms](/users/max-harms) on [Empowerment, corrigibility, etc. are simple abstractions (of a messed-up ontology)](/api/post/empowerment-corrigibility-etc-are-simple-abstractions-of-a) * 2026-05-12 20:37:31Z * Karma: 18 * Total votes: 11 * Comment URL (Markdown): [/api/post/empowerment-corrigibility-etc-are-simple-abstractions-of-a/comments/meqL6TXfhTGDGKzAn](/api/post/empowerment-corrigibility-etc-are-simple-abstractions-of-a/comments/meqL6TXfhTGDGKzAn) * Comment URL (HTML): [/posts/vzHtHHBJoKATi5SeK/empowerment-corrigibility-etc-are-simple-abstractions-of-a/comment/meqL6TXfhTGDGKzAn](/posts/vzHtHHBJoKATi5SeK/empowerment-corrigibility-etc-are-simple-abstractions-of-a/comment/meqL6TXfhTGDGKzAn) I think you're correctly identifying important issues and cracks in the standard ontology, but I think you're throwing out too much baby in an effort to get rid of bathwater. For example, I do not think it's obvious that "Just like vitalistic force, 'wanting' is conceptualized as being acausal, i.e. an intrinsic property of an entity with no upstream cause." In control theory, we can say that a system controls for a thing based on a small collection of mathematical relationships -- pressuring an error signal towards zero. While the concept of wanting is overloaded and more complex, I think it makes sense to recognize that "X is controlling for Y" is a valid underpinning that has no vitalistic magic. We can ask what led X to control for Y, or how X controls for Y in terms that are closer to the underlying physics; there's nothing acausal or intrinsic (except for the definitions, I suppose). ### Comment by [Max Harms](/users/max-harms) on [Empowerment, corrigibility, etc. are simple abstractions (of a messed-up ontology)](/api/post/empowerment-corrigibility-etc-are-simple-abstractions-of-a) * 2026-05-12 20:17:38Z * Karma: 3 * Total votes: 2 * Comment URL (Markdown): [/api/post/empowerment-corrigibility-etc-are-simple-abstractions-of-a/comments/nDgan3jeva48PiRsk](/api/post/empowerment-corrigibility-etc-are-simple-abstractions-of-a/comments/nDgan3jeva48PiRsk) * Comment URL (HTML): [/posts/vzHtHHBJoKATi5SeK/empowerment-corrigibility-etc-are-simple-abstractions-of-a/comment/nDgan3jeva48PiRsk](/posts/vzHtHHBJoKATi5SeK/empowerment-corrigibility-etc-are-simple-abstractions-of-a/comment/nDgan3jeva48PiRsk) Oh, uh. Whoops. Forgot to switch to my work account. ### Comment by [Max Harms](/users/max-harms) on [1\. The CAST Strategy](/api/post/1-the-cast-strategy) * 2026-01-08 21:49:17Z * Karma: 5 * Total votes: 3 * Comment URL (Markdown): [/api/post/1-the-cast-strategy/comments/Fn9LtaxcPbLrXxnnE](/api/post/1-the-cast-strategy/comments/Fn9LtaxcPbLrXxnnE) * Comment URL (HTML): [/posts/3HMh7ES4ACpeDKtsW/1-the-cast-strategy/comment/Fn9LtaxcPbLrXxnnE](/posts/3HMh7ES4ACpeDKtsW/1-the-cast-strategy/comment/Fn9LtaxcPbLrXxnnE) Thanks. I'll put most of my thoughts in a comment on your post, but I guess I want to say here that the issues you raise are adjacent to the reasons I listed "write a guide" as the second option, rather than the first (i.e. surveillance + ban). We need plans that we can be confident in even while grappling with how lost we are on the ethical front. ### Navigation * [Front page](https://www.alignmentforum.org/api/home) * [Markdown API documentation](https://www.alignmentforum.org/api/SKILL.md)