In the previous post, I presented my theory of change for why value generalisation is vital for AI alignment. Here I'll add the practical part of the argument: given those facts, why explicitly try to do value generalisation, what are the dangers of the approach, how should it be done, and how do we mitigate the risks?
The formal theory of change is down below, but I'll put a collapsible version here, to make references easier:
Theory of Change
Inputs / activities: investment/grants, a small research team, small-medium compute resources.
Outputs:
Solutions to the three components of value generalisation (see here) in academically rigorous forms, demonstrations of these solutions on toy problems, public benchmarks, and new benchmarks. If going commercial, applications of this value generalisation to saleable products.
The first component is recognising that the AI is off-distribution in a value-relevant way.
The second component is establishing what features should be used to reach a decision while off-distribution this way.
The third component is to reach a good decision while off distribution, and to loop through the process again to critique and improve the decision, including using outside feedback.
Combination of these approaches in a successful prealigned learning model.
Development of value generalisation alignment methods for AIs.
Making these methods available to AI safety researchers.
Prealigned models being generally available to safely empower users and keep humanity on track for a positive outcome.
If the approach fails, the failed attempts and the partial successes will be made available to others.
Impact:
If failed:
A seemingly promising approach crossed off the list of possible paths to alignment.
A better understanding of the issues around the "value generalisation" framing.
If intermediate outcome:
A tool or collection of useful tools or models that can improve some AI alignment techniques.
If successful:
AI alignment, or a big step forwards towards it.
An increase in AI capabilities.
Explicit targeting of value generalisation
Claim : explicit value generalisation is much more useful than implicit value generalisation.
Fundamentally, there's a difference between an AI that knows what course humans would consider the best one, and one that actually follows that course. Having the AI capable of value generalisation inside its own mind is not useful, unless we can incorporate that generalisation into its goals. And explicit generalisation allows that.
This leads to the crucial and unfortunate claim:
Claim : value generalisation will aid in empirical generalisation. But empirical generalisation will likely not aid value generalisation.
Empirical generalisation is the ability to update features usefully across model splinterings, in ways that preserve or improve the ability of the modeller to effectively influence the world.
Claim derives in part from claim : since useful value generalisation is the explicit kind, implicit empirical generalisation is not likely to usefully help. It also derives from the fact that empirical generalisation can afford to discard its features if necessary: we refined concepts like heat while discarding vitalism. But we can't just discard the "suffering" feature in "avoid human suffering"; it has to be redefined and extended. So value generalisation is strictly harder.
Somewhat connected to that is the fact that "in the limit" of infinite computation and infinite observation, empirical generalisation is doable, but value generalisation involves moral choices that don't come free from mere observations.
More details on the dis-equivalence between implicit or explicit empirical generalisation, and explicit value generalisation
Assume we have , a set of environments, with a probability distribution over . An agent uses features to model the set of environments; these can be seen as numerical functions on the set of environments, so environment is modelled by the value , and there is a probability distribution over the values of the features. These features will include the agent's actions, allowing it to plot a policy.
For simplicity, assume the features are Boolean valued[1], so a world is mapped, via the , to an element of , which is also written as .
Then give a utility function , we can say that is useful for -maximising if there exists a utility function such that policies that optimise or fail to optimise , will also optimise or fail to optimise . Thus agent can stay in their modelling space and plan and act there.
Typically, the agent doesn't start with access to ground reality, so they don't start with , but with a . Note that just because a is defined over features, doesn't mean that the features are useful for maximising that function. The score in Pacman is enough of a feature to define an objective, but not enough to play the game.
A model splintering is a change in and/or , replacing with (a model splintering might be as simple as the agent realising they have extra options - such as in reward tampering or wire-heading - or it might be a complete change of their view of reality). Note that model splinterings are not things the agent directly observes - the need for a model splintering is inferred from anomalous-seeming observations.
Empirical feature generalisation is an algorithm that takes in the agent's internal state (which corresponds to ) and the agent's history . The then moves the agent to an internal state (which corresponds to new features ), such that the new features are useful for maximising the original (or ), given the evidence has provided about model splintering.
Explicit feature generalisation does the process in terms of the features themselves: maps and to the new . Technically, neither nor define each other; but it is generally much easier to construct a black-box method from an explicit method than the opposite. And, in the limit, might be constructible by brute force analysis.
Value generalisation instead uses an that plays the same role as , except that it also has to map to a new . Empirical generalisation (or just exploration) might establish that spinning the boat endlessly gives maximal points; value generalisation also has to figure out that this is not a desired generalisation of the original goal. For a similar reason, while can discard a feature as no longer useful, or simplify it to the extreme, can't discard or over-simplify a value-relevant feature; indeed, these features are often expected to grow more complex over time (e.g. the definition of sentient being).
But why "also"? Could , or the explicit version , not simply focus on the value generalisation and ignore the empirical changes? The first problem is that, in order to generalise , the agent will have to figure out a lot about the features of the environment that are correlated or not correlated with and how they relate to each other (and how that relationship might be changeable).
For instance, human smiling is correlated with the feeling of happiness, the release of certain hormones, physiological changes in the brain, reported happiness, and so on. So to generalise "smiling" to "true happiness" and beyond, the agent has to learn a lot about the structure of the world and what features best describe it. It also has to learn how to change the various potential : if a given candidate is suddenly very easy to optimise, that's a sign that it might be a Goodhart proxy.
The second problem is that could be replaced with any instrumental goal , and any change of .
So it seems that , which we need for effective and checkable value generalisation, must be able to do a deep analysis of any instrumental goal over any model splintering, including figuring out ways to optimise the original and the new instrumental goals. This seems to feed straight into empirical generalisation.
I will ruthlessly search for that don't give powerful , but I am not optimistic.
This asymmetry leads to the unfortunate result that:
Claim : a value generalisation project will increase AI capabilities.
Claim : strong generalisation is a uniquely human ability, that current LLMs don't have and will likely not develop; nor will any similar model develop strong generalisation either.
However, this claim is not crucial to the approach. If is wrong, claim becomes less worrying (AIs will develop the ability anyway) but the research becomes more urgent: we need to get a decent start on explicit value generalisation before AIs get good at empirical generalisation.
The weaker claim is:
Claim : studying how humans do strong generalisation is a fruitful way of analysing the skill, at least initially.
The arguments in the second value generalisation post apply also for claim . Even if one is of the claim that "LLMs will always fail at strong generalisation", there is certainly evidence that humans are doing something very different (and much more data-efficient) when we generalise.
The question of timing
Given that value generalisation is essential to alignment, but also could lead to capability increase, the question is: should we research it now (early) or later (late)?
I believe:
Claim : value generalisation should be researched early.
The arguments for this are that it's better to have a capability increase while the models are weaker and more controllable, and early value generalisation can more easily be integrated into models from the get-go rather than retro-fitting at a later date. We want a minimal capability overhang that empirical generalisation can unleash. And, of course, we want to avoid the scenarios where the research arrives too late.
The corporate argument
Given the above, I propose creating either a research program or a commercial entity to find useable solutions to value generalisation. But I will claim:
Claim : if value generalisation is indeed revolutionary but does lead to dangerous capability increases, then the commercial route weakly dominates the research route.
The main argument for claim comes from the question: assume that value generalisation is solved or partially solved, then what? We have a powerful alignment technology that is essential for alignment, but also a powerful capability technology, in a way that can't be separated. What do we do with it?
Well, we'd probably want to hand it over to some trustworthy entity (some have suggested a "CERN for AI") to implement the alignment part at some point.
But that is easier to achieve for a corporation than for a research program. A corporation is much better placed to keep its research private or patent-protected and concealed from the world. It can sell or license the research to a trustworthy entity, or sell it to a semi-trustworthy entity with conditions. And it can also choose to be the trustworthy entity and start implementing alignment itself. Indeed:
Claim : if the capability boost from generalisation is inevitable, a corporation can mix the capability increase with the alignment increase (such as prealigned AIs) so that the first generalising AIs are aligned, and the initial income from generalising AIs flows to value aligned entities.
I've mentioned the advantages of the corporate route, but what of the advantages of the academic or research route - such as openness allowing more scrutiny and feedback, getting more trust from the safety community? Well, the research is potentially dangerous, so we can't expect to have it publicly or semi-publicly available. Thus:
Claim : the potential need for secrecy reduces the standard advantages from the academic or research institution route.
Armouring the corporate weak points
Of course, the corporate path has its weaknesses, mainly revolving around the profit motive. Investors, the legal system, management, and employees will all want a company to cash in on legal innovative ideas, even if the ideas are potentially dangerous. So a first step would be:
Design : the corporation should have an independent AI ethics board, with the power to block the use of IP it deems dangerous. To facilitate this, the AI ethics board should be the owner of the IP.
That solves the problem in the formal sense, which means that it doesn't really solve it. Additional measures would be:
Design : the employees and management should be value aligned with AI safety.
Design : the investors should be value-aligned with AI safety.
That also helps; also not enough. The situation will never be "do I kill everyone with certainty to make $10,000 more this year"? The situation will be more like "when I feel the ethics board is being fussy and unreasonable and overly cautious, do I nevertheless bow down to their irrational decrees that will cost me a lot of my expected income that I have worked so hard on and so earned, and also give up my possibility of improving the world for the better"? And value-aligned investors remain investors: they expect to make a profit (a not-unreasonable demand, which the legal system backs up).
I don't expect myself to be immune to that pressure and those arguments. And it's not just a question of holding firm; as OpenAI's experience demonstrates:
Claim : if the employees are ready to jump ship to a partner organisation that can continue the research with a minimum of fuss, the AI ethics board's power is theoretical rather than real.
So I've been considering ways to remove the stark tension. One design is:
Design : the AI ethics board will have the power to implement a "pivot to immediate profit". The long-range dangerous IP will be removed from the company, and the company will turn to making money from all the partial ideas and designs that it has made and rejected (as not being alignment relevant) along the way. The board will, of course, vet these ideas for danger, but immediate short-term profitability is less likely to overlap with truly dangerous IP.
Arguably, the pivot to immediate profit would be more profitable, in expectation, than speculative long term IP. This would relieve much of the commercial pressure. And if the long term IP was truly world-improving, the ethics board would find a way to get it carefully deployed, so world-improving impulses are preserved (if the research is merely dangerous, the ethics board would bury it, and good riddance).
Plan, milestones, and assumption checks
Let's gather all the previous together to put it all in one plan:
Inputs / activities: investment/grants, a small research team, small-medium compute resources.
Outputs:
Solutions to the three components of value generalisation (see here) in academically rigorous forms, demonstrations of these solutions on toy problems, public benchmarks, and new benchmarks. If going commercial, applications of this value generalisation to saleable products.
The first component is recognising that the AI is off-distribution in a value-relevant way.
The second component is establishing what features should be used to reach a decision while off-distribution this way.
The third component is to reach a good decision while off distribution, and to loop through the process again to critique and improve the decision, including using outside feedback.
Combination of these approaches in a successful prealigned learning model.
Development of value generalisation alignment methods for AIs.
Making these methods available to AI safety researchers.
Prealigned models being generally available to safely empower users and keep humanity on track for a positive outcome.
If the approach fails, the failed attempts and the partial successes will be made available to others.
Impact:
If failed:
A seemingly promising approach crossed off the list of possible paths to alignment.
A better understanding of the issues around the "value generalisation" framing.
If intermediate outcome:
A tool or collection of useful tools or models that can improve some AI alignment techniques.
If successful:
AI alignment, or a big step forwards towards it.
An increase in AI capabilities.
Ongoing and initial assessments
This theory of change has been a bit light on specific probability estimates for different assumptions, and for the overall program.
The reason is that, for this approach, the proof of the pudding is in the eating. Now, I've presented solid theoretical arguments for the necessity and usefulness of value generalisation, I can point to my own research track record and intermediate successfulresults and a plan for the R&D path and how it builds on human generalisation abilities.
I feel the case is compelling, but it rests ultimately on a judgement call: that it makes sense to group all these problems under the heading "value generalisation", to see it all as a single specific goal, and to aim to tackle it directly.
That judgement call may be valid (I certainly feel it is!) and is the true crux in this theory of change. And the best way of figuring out its truth or falsity is to... attempt the project and see what happens, see whether the framework gives swift success or falls apart into disparate disconnected problems.
To check on that, we'll need intermediate benchmarks and assessments. The initial phase of the program (carried out in part while the setup is happening) will include establishing key benchmarks for each subsequent step.
We'll be using public benchmarks for these purposes, though they need to be used with care[2]. A lot of benchmarks get saturated quite easily by models that don't show the true performance the benchmark was supposed to measure. We need algorithms that solve benchmarks via value generalisation, not via any intermediate incomplete shortcuts. The benchmark design can do part of the work, but controlling the information and methods the algorithm can use is also crucial.
Conclusion: value generalisation is a useful path for AI alignment
To summarise:
Most AI alignment failures are value generalisation failures.
Non-decomposability: alignment likely can't be decomposed into smaller simpler parts.
Thus value generalisation is necessary for alignment, and will make alignment a lot easier.
Explicitly targeting value generalisation is necessary.
Unfortunately, value generalisation aids empirical generalisation (a capability) while the converse is not true.
A corporate structure weakly dominates a research structure for solving value generalisation and for using the results as safely as possible.
Good design and a good AI ethics board can mitigate the vulnerabilities of the corporate route.
There is a plan to solve value alignment step by step, focusing first on how agents identify they are out of distribution, then what features to use for deciding in those circumstances, and finally what decision to make and how to assess and improve that decision.
So, let's go and solve this little alignment problem, aye? ^_^
Take the "wolf vs husky" image classification problem, where the wolf images were all taken on snow, causing the classifier to misclassify any light-background image as a wolf.
But a well-trained image recognition model could likely solve both of these benchmarks, just by having enough wolf/dog/tank/forest data that it can resolve these images anyway. That's not out-of-distribution learning, that's increasing the training data until the benchmarks are in-distribution.
In the previous post, I presented my theory of change for why value generalisation is vital for AI alignment. Here I'll add the practical part of the argument: given those facts, why explicitly try to do value generalisation, what are the dangers of the approach, how should it be done, and how do we mitigate the risks?
The formal theory of change is down below, but I'll put a collapsible version here, to make references easier:
Theory of Change
Explicit targeting of value generalisation
Fundamentally, there's a difference between an AI that knows what course humans would consider the best one, and one that actually follows that course. Having the AI capable of value generalisation inside its own mind is not useful, unless we can incorporate that generalisation into its goals. And explicit generalisation allows that.
This leads to the crucial and unfortunate claim:
Empirical generalisation is the ability to update features usefully across model splinterings, in ways that preserve or improve the ability of the modeller to effectively influence the world.
Claim derives in part from claim : since useful value generalisation is the explicit kind, implicit empirical generalisation is not likely to usefully help. It also derives from the fact that empirical generalisation can afford to discard its features if necessary: we refined concepts like heat while discarding vitalism. But we can't just discard the "suffering" feature in "avoid human suffering"; it has to be redefined and extended. So value generalisation is strictly harder.
Somewhat connected to that is the fact that "in the limit" of infinite computation and infinite observation, empirical generalisation is doable, but value generalisation involves moral choices that don't come free from mere observations.
More details on the dis-equivalence between implicit or explicit empirical generalisation, and explicit value generalisation
Assume we have , a set of environments, with a probability distribution over . An agent uses features to model the set of environments; these can be seen as numerical functions on the set of environments, so environment is modelled by the value , and there is a probability distribution over the values of the features. These features will include the agent's actions, allowing it to plot a policy.
For simplicity, assume the features are Boolean valued[1], so a world is mapped, via the , to an element of , which is also written as .
Then give a utility function , we can say that is useful for -maximising if there exists a utility function such that policies that optimise or fail to optimise , will also optimise or fail to optimise . Thus agent can stay in their modelling space and plan and act there.
Typically, the agent doesn't start with access to ground reality, so they don't start with , but with a . Note that just because a is defined over features, doesn't mean that the features are useful for maximising that function. The score in Pacman is enough of a feature to define an objective, but not enough to play the game.
A model splintering is a change in and/or , replacing with (a model splintering might be as simple as the agent realising they have extra options - such as in reward tampering or wire-heading - or it might be a complete change of their view of reality). Note that model splinterings are not things the agent directly observes - the need for a model splintering is inferred from anomalous-seeming observations.
Empirical feature generalisation is an algorithm that takes in the agent's internal state (which corresponds to ) and the agent's history . The then moves the agent to an internal state (which corresponds to new features ), such that the new features are useful for maximising the original (or ), given the evidence has provided about model splintering.
Explicit feature generalisation does the process in terms of the features themselves: maps and to the new . Technically, neither nor define each other; but it is generally much easier to construct a black-box method from an explicit method than the opposite. And, in the limit, might be constructible by brute force analysis.
Value generalisation instead uses an that plays the same role as , except that it also has to map to a new . Empirical generalisation (or just exploration) might establish that spinning the boat endlessly gives maximal points; value generalisation also has to figure out that this is not a desired generalisation of the original goal. For a similar reason, while can discard a feature as no longer useful, or simplify it to the extreme, can't discard or over-simplify a value-relevant feature; indeed, these features are often expected to grow more complex over time (e.g. the definition of sentient being).
But why "also"? Could , or the explicit version , not simply focus on the value generalisation and ignore the empirical changes? The first problem is that, in order to generalise , the agent will have to figure out a lot about the features of the environment that are correlated or not correlated with and how they relate to each other (and how that relationship might be changeable).
For instance, human smiling is correlated with the feeling of happiness, the release of certain hormones, physiological changes in the brain, reported happiness, and so on. So to generalise "smiling" to "true happiness" and beyond, the agent has to learn a lot about the structure of the world and what features best describe it. It also has to learn how to change the various potential : if a given candidate is suddenly very easy to optimise, that's a sign that it might be a Goodhart proxy.
The second problem is that could be replaced with any instrumental goal , and any change of .
So it seems that , which we need for effective and checkable value generalisation, must be able to do a deep analysis of any instrumental goal over any model splintering, including figuring out ways to optimise the original and the new instrumental goals. This seems to feed straight into empirical generalisation.
I will ruthlessly search for that don't give powerful , but I am not optimistic.
This asymmetry leads to the unfortunate result that:
How to achieve value generalisation
I've previously made the claim that:
However, this claim is not crucial to the approach. If is wrong, claim becomes less worrying (AIs will develop the ability anyway) but the research becomes more urgent: we need to get a decent start on explicit value generalisation before AIs get good at empirical generalisation.
The weaker claim is:
The arguments in the second value generalisation post apply also for claim . Even if one is of the claim that "LLMs will always fail at strong generalisation", there is certainly evidence that humans are doing something very different (and much more data-efficient) when we generalise.
The question of timing
Given that value generalisation is essential to alignment, but also could lead to capability increase, the question is: should we research it now (early) or later (late)?
I believe:
The arguments for this are that it's better to have a capability increase while the models are weaker and more controllable, and early value generalisation can more easily be integrated into models from the get-go rather than retro-fitting at a later date. We want a minimal capability overhang that empirical generalisation can unleash. And, of course, we want to avoid the scenarios where the research arrives too late.
The corporate argument
Given the above, I propose creating either a research program or a commercial entity to find useable solutions to value generalisation. But I will claim:
The main argument for claim comes from the question: assume that value generalisation is solved or partially solved, then what? We have a powerful alignment technology that is essential for alignment, but also a powerful capability technology, in a way that can't be separated. What do we do with it?
Well, we'd probably want to hand it over to some trustworthy entity (some have suggested a "CERN for AI") to implement the alignment part at some point.
But that is easier to achieve for a corporation than for a research program. A corporation is much better placed to keep its research private or patent-protected and concealed from the world. It can sell or license the research to a trustworthy entity, or sell it to a semi-trustworthy entity with conditions. And it can also choose to be the trustworthy entity and start implementing alignment itself. Indeed:
I've mentioned the advantages of the corporate route, but what of the advantages of the academic or research route - such as openness allowing more scrutiny and feedback, getting more trust from the safety community? Well, the research is potentially dangerous, so we can't expect to have it publicly or semi-publicly available. Thus:
Armouring the corporate weak points
Of course, the corporate path has its weaknesses, mainly revolving around the profit motive. Investors, the legal system, management, and employees will all want a company to cash in on legal innovative ideas, even if the ideas are potentially dangerous. So a first step would be:
That solves the problem in the formal sense, which means that it doesn't really solve it. Additional measures would be:
That also helps; also not enough. The situation will never be "do I kill everyone with certainty to make $10,000 more this year"? The situation will be more like "when I feel the ethics board is being fussy and unreasonable and overly cautious, do I nevertheless bow down to their irrational decrees that will cost me a lot of my expected income that I have worked so hard on and so earned, and also give up my possibility of improving the world for the better"? And value-aligned investors remain investors: they expect to make a profit (a not-unreasonable demand, which the legal system backs up).
I don't expect myself to be immune to that pressure and those arguments. And it's not just a question of holding firm; as OpenAI's experience demonstrates:
So I've been considering ways to remove the stark tension. One design is:
Arguably, the pivot to immediate profit would be more profitable, in expectation, than speculative long term IP. This would relieve much of the commercial pressure. And if the long term IP was truly world-improving, the ethics board would find a way to get it carefully deployed, so world-improving impulses are preserved (if the research is merely dangerous, the ethics board would bury it, and good riddance).
Plan, milestones, and assumption checks
Let's gather all the previous together to put it all in one plan:
Ongoing and initial assessments
This theory of change has been a bit light on specific probability estimates for different assumptions, and for the overall program.
The reason is that, for this approach, the proof of the pudding is in the eating. Now, I've presented solid theoretical arguments for the necessity and usefulness of value generalisation, I can point to my own research track record and intermediate successful results and a plan for the R&D path and how it builds on human generalisation abilities.
I feel the case is compelling, but it rests ultimately on a judgement call: that it makes sense to group all these problems under the heading "value generalisation", to see it all as a single specific goal, and to aim to tackle it directly.
That judgement call may be valid (I certainly feel it is!) and is the true crux in this theory of change. And the best way of figuring out its truth or falsity is to... attempt the project and see what happens, see whether the framework gives swift success or falls apart into disparate disconnected problems.
To check on that, we'll need intermediate benchmarks and assessments. The initial phase of the program (carried out in part while the setup is happening) will include establishing key benchmarks for each subsequent step.
We'll be using public benchmarks for these purposes, though they need to be used with care[2]. A lot of benchmarks get saturated quite easily by models that don't show the true performance the benchmark was supposed to measure. We need algorithms that solve benchmarks via value generalisation, not via any intermediate incomplete shortcuts. The benchmark design can do part of the work, but controlling the information and methods the algorithm can use is also crucial.
Conclusion: value generalisation is a useful path for AI alignment
To summarise:
So, let's go and solve this little alignment problem, aye? ^_^
Any function taking values can be represented by Boolean functions.
Take the "wolf vs husky" image classification problem, where the wolf images were all taken on snow, causing the classifier to misclassify any light-background image as a wolf.
We had a similar benchmark, "tanks vs forests" where the tank images were taken on a cloudy day (a recreation of an apocryphal but traditional tale in machine learning). The ACE algorithm successfully disambiguated "darkness" from "tankiness".
But a well-trained image recognition model could likely solve both of these benchmarks, just by having enough wolf/dog/tank/forest data that it can resolve these images anyway. That's not out-of-distribution learning, that's increasing the training data until the benchmarks are in-distribution.