# User: Scott Garrabrant Profile URL (HTML): [/users/scott-garrabrant](/users/scott-garrabrant) Profile URL (Markdown): [/api/user/scott-garrabrant](/api/user/scott-garrabrant) * Karma: 8908 * Alignment Forum karma: 1614 * Posts: 74 * Comments: 416 * Member since: 2017-09-22 02:21:16Z Bio --- *No bio.* Top Posts --------- ### [Embedded Agents](/api/post/embedded-agents) By [abramdemski](/users/abramdemski) with [Scott Garrabrant](/users/scott-garrabrant) 2018-10-29 19:53:02Z * Karma: 252 * Curated * Tags: [Research Agendas](/w/research-agendas), [Embedded Agency](/w/embedded-agency), [AI](/w/ai) (Frontpage) Read more: [/api/post/embedded-agents](/api/post/embedded-agents) ### [Embedded Agency (full-text version)](/api/post/embedded-agency-full-text-version) By [Scott Garrabrant](/users/scott-garrabrant) with [abramdemski](/users/abramdemski) 2018-11-15 19:49:29Z * Karma: 235 * Curated * Tags: [Embedded Agency](/w/embedded-agency), [Agent Foundations](/w/agent-foundations), [Mesa-Optimization](/w/mesa-optimization), [Decision theory](/w/decision-theory), [Research Agendas](/w/research-agendas), [Goodhart's Law](/w/goodhart-s-law), [Robust Agents](/w/robust-agents), [Spurious Counterfactuals](/w/spurious-counterfactuals), [Subagents](/w/subagents), [Boundaries / Membranes \[technical\]](/w/boundaries-membranes-technical), [AI](/w/ai) (Frontpage) Read more: [/api/post/embedded-agency-full-text-version](/api/post/embedded-agency-full-text-version) ### [Goodhart Taxonomy](/api/post/goodhart-taxonomy) By [Scott Garrabrant](/users/scott-garrabrant) 2017-12-30 16:38:39Z * Karma: 229 * Curated * Tags: [Goodhart's Law](/w/goodhart-s-law), [AI](/w/ai) (Frontpage) Read more: [/api/post/goodhart-taxonomy](/api/post/goodhart-taxonomy) Recent Posts ------------ ### [Concave Utility Question](/api/post/concave-utility-question) By [Scott Garrabrant](/users/scott-garrabrant) 2023-04-15 00:14:58Z * Karma: 55 * Tags: [AI](/w/ai) (Frontpage) Read more: [/api/post/concave-utility-question](/api/post/concave-utility-question) ### [Counterfactability](/api/post/counterfactability) By [Scott Garrabrant](/users/scott-garrabrant) 2022-11-07 05:39:05Z * Karma: 40 * Tags: [Finite Factored Sets](/w/finite-factored-sets), [AI](/w/ai), [World Modeling](/w/world-modeling) (Frontpage) Read more: [/api/post/counterfactability](/api/post/counterfactability) ### [Boundaries vs Frames](/api/post/boundaries-vs-frames) By [Scott Garrabrant](/users/scott-garrabrant) 2022-10-31 15:14:37Z * Karma: 58 * Tags: [Boundaries / Membranes \[technical\]](/w/boundaries-membranes-technical), [AI](/w/ai) (Frontpage) Read more: [/api/post/boundaries-vs-frames](/api/post/boundaries-vs-frames) ### [Cartesian Frames and Factored Sets on ArXiv](/api/post/cartesian-frames-and-factored-sets-on-arxiv) By [Scott Garrabrant](/users/scott-garrabrant) 2021-09-24 04:58:54Z * Karma: 38 * Tags: [Finite Factored Sets](/w/finite-factored-sets), [World Modeling](/w/world-modeling) (Frontpage) Read more: [/api/post/cartesian-frames-and-factored-sets-on-arxiv](/api/post/cartesian-frames-and-factored-sets-on-arxiv) ### [Finite Factored Sets: Applications](/api/post/finite-factored-sets-applications) By [Scott Garrabrant](/users/scott-garrabrant) 2021-08-31 21:19:03Z * Karma: 34 * Tags: [Finite Factored Sets](/w/finite-factored-sets), [AI](/w/ai), [World Modeling](/w/world-modeling) (Frontpage) Read more: [/api/post/finite-factored-sets-applications](/api/post/finite-factored-sets-applications) ### [Finite Factored Sets: Inferring Time](/api/post/finite-factored-sets-inferring-time) By [Scott Garrabrant](/users/scott-garrabrant) 2021-08-31 21:18:36Z * Karma: 23 * Tags: [Finite Factored Sets](/w/finite-factored-sets), [AI](/w/ai) (Frontpage) Read more: [/api/post/finite-factored-sets-inferring-time](/api/post/finite-factored-sets-inferring-time) ### [Finite Factored Sets: Polynomials and Probability](/api/post/finite-factored-sets-polynomials-and-probability) By [Scott Garrabrant](/users/scott-garrabrant) 2021-08-17 21:53:03Z * Karma: 21 * Tags: [Finite Factored Sets](/w/finite-factored-sets), [AI](/w/ai), [Rationality](/w/rationality) (Frontpage) Read more: [/api/post/finite-factored-sets-polynomials-and-probability](/api/post/finite-factored-sets-polynomials-and-probability) ### [Finite Factored Sets: Conditional Orthogonality](/api/post/finite-factored-sets-conditional-orthogonality) By [Scott Garrabrant](/users/scott-garrabrant) 2021-07-09 06:01:46Z * Karma: 29 * Tags: [Finite Factored Sets](/w/finite-factored-sets), [Abstraction](/w/abstraction), [Causality](/w/causality), [AI](/w/ai), [Rationality](/w/rationality) (Frontpage) Read more: [/api/post/finite-factored-sets-conditional-orthogonality](/api/post/finite-factored-sets-conditional-orthogonality) ### [Finite Factored Sets: LW transcript with running commentary](/api/post/finite-factored-sets-lw-transcript-with-running-commentary) By [Rob Bensinger](/users/robbbb) with [Scott Garrabrant](/users/scott-garrabrant) 2021-06-27 16:02:06Z * Karma: 30 * Tags: [Finite Factored Sets](/w/finite-factored-sets), [AI](/w/ai) (Frontpage) Read more: [/api/post/finite-factored-sets-lw-transcript-with-running-commentary](/api/post/finite-factored-sets-lw-transcript-with-running-commentary) ### [Finite Factored Sets: Orthogonality and Time](/api/post/finite-factored-sets-orthogonality-and-time) By [Scott Garrabrant](/users/scott-garrabrant) 2021-06-10 01:22:34Z * Karma: 34 * Tags: [Finite Factored Sets](/w/finite-factored-sets) (Frontpage) Read more: [/api/post/finite-factored-sets-orthogonality-and-time](/api/post/finite-factored-sets-orthogonality-and-time) Recent Comments --------------- ### Comment by [Scott Garrabrant](/users/scott-garrabrant) on [Infrafunctions and Robust Optimization](/api/post/infrafunctions-and-robust-optimization) * 2023-05-02 01:13:54Z * Karma: 23 * Total votes: 5 * Comment URL (Markdown): [/api/post/infrafunctions-and-robust-optimization/comments/t4FmopzcQqbm8n3Yw](/api/post/infrafunctions-and-robust-optimization/comments/t4FmopzcQqbm8n3Yw) * Comment URL (HTML): [/posts/d96dDEYMfnN2St3Bj/infrafunctions-and-robust-optimization/comment/t4FmopzcQqbm8n3Yw](/posts/d96dDEYMfnN2St3Bj/infrafunctions-and-robust-optimization/comment/t4FmopzcQqbm8n3Yw) Here are the most interesting things about these objects to me that I think this post does not capture.  Given a distribution over non-negative non-identically-zero infrafunctions, up to a positive scalar multiple, the pointwise geometric expectation exists, and is an infra function (up to a positive scalar multiple). (I am not going to give all the math and be careful here, but hopefully this comment will provide enough of a pointer if someone wants to investigate this.) This is a bit of a miracle. Compare this with arithmetic expectation of utility functions. This is not always well defined. For example, if you have a sequence of utility functions U_n, each with weight 2^{-n}, but which alternate in which of two outcomes they prefer, and each utility function gets an internal weighting to cancel out their small weight an then some, the expected utility will not exist. There will be a series of larger and larger utility monsters canceling each other out, and the limit will not exist. You could fix this requiring your utility functions are bounded, as is standard for dealing with utility monsters, but it is really interesting that in the case of infra functions and geometric expectation, you don't have to. If you try to do a similar trick with infra functions, up to a positive scalar multiple, geometric expectation will go to infinity, but you can renormalize everything since you are only working up to a scalar multiple, to make things well defined. We needed the geometric expectation to only be working up to a scalar multiple, and you cant expect a utility function if you take a geometric expectation of utility functions. (but you do get an infrafunction!) If you start with utility functions, and then merge them geometrically, the resulting infrafunction will be maximized at the Nash bargaining solution, but the entire infrafunction can be thought of as an extended preference over lotteries of the pair of utility functions, where as Nash bargaining only told you the maximum. In this way geometric merging of infrafunctions is starting with an input more general than the utility functions of Nash bargaining, and giving an output more structured than the output of Nash bargaining, and so can be thought of as a way of making Nash bargaining more compositional. (Since the input and output are now the same type, so you can stack them on top of each other.) For these two reasons (utility monster resistance and extending Nash bargaining), I am very interested in the mathematical object that is non-negative non-identically-zero infrafunctions defined only up to a positive scalar multiple, and more specifically, I am interested in the set of such functions as a *convex* set where mixing is interpreted as pointwise geometric expectation. ### Comment by [Scott Garrabrant](/users/scott-garrabrant) on [Infrafunctions and Robust Optimization](/api/post/infrafunctions-and-robust-optimization) * 2023-05-02 00:54:42Z * Karma: 10 * Total votes: 4 * Comment URL (Markdown): [/api/post/infrafunctions-and-robust-optimization/comments/Ta8roBMRseXuFAbAC](/api/post/infrafunctions-and-robust-optimization/comments/Ta8roBMRseXuFAbAC) * Comment URL (HTML): [/posts/d96dDEYMfnN2St3Bj/infrafunctions-and-robust-optimization/comment/Ta8roBMRseXuFAbAC](/posts/d96dDEYMfnN2St3Bj/infrafunctions-and-robust-optimization/comment/Ta8roBMRseXuFAbAC) I have been thinking about this [same mathematical object](/api/post/uJnR4YmG5Kq9FfTey) (although with a different orientation/motivation) as where I want to go with a weaker replacement for utility functions. I get the impression that for Diffractor/Vanessa, the heart of a concave-value-function-on-lotteries is that it represents the worst case utility over some set of possible utility functions. For me, on the other hand, a concave value function represents the capacity for compromise -- if I get at least half the good if I get what I want with 50% probability, then I have the capacity to merge/compromise with others using tools like Nash bargaining.  This brings us to the same mathematical object, but it feels like I am using the definition of convex set related to the line segment connecting any two points in the set is also in the set, where Diffractor/Vanessa is using the definition of convex set related to being an intersection of half planes.  I think this pattern where I am more interested in merging, and Diffractor and Vanessa are more interested in guarantees, but we end up looking at the same math is a pattern, and I think the dual definitions of convex set in part explains (or at least rhymes with) this pattern. ### Comment by [Scott Garrabrant](/users/scott-garrabrant) on [Concave Utility Question](/api/post/concave-utility-question) * 2023-04-17 07:59:18Z * Karma: 4 * Total votes: 2 * Comment URL (Markdown): [/api/post/concave-utility-question/comments/dLTwHGJ4hhLEMfogH](/api/post/concave-utility-question/comments/dLTwHGJ4hhLEMfogH) * Comment URL (HTML): [/posts/uJnR4YmG5Kq9FfTey/concave-utility-question/comment/dLTwHGJ4hhLEMfogH](/posts/uJnR4YmG5Kq9FfTey/concave-utility-question/comment/dLTwHGJ4hhLEMfogH) Then it is equivalent to the thing I call B2 in edit 2 in the post (Assuming A1-A3). In this case, your modified B2 is my B2, and your B3 is my A4, which follows from A5 assuming A1-A3 and B2, so your suspicion that these imply C4 is stronger than my Q6, which is false, as I argue [here](/api/post/uJnR4YmG5Kq9FfTey?commentId=47igAoaWde3Hyfuq7). However, without A5, it is actually much easier to see that this doesn't work. The counterexample [here](/api/post/uJnR4YmG5Kq9FfTey?commentId=LKu36ngqfu3CGHjWv) satisfies my A1-A3, your weaker version of B2, your B3, and violates C4. ### Comment by [Scott Garrabrant](/users/scott-garrabrant) on [Concave Utility Question](/api/post/concave-utility-question) * 2023-04-16 18:49:42Z * Karma: 2 * Total votes: 1 * Comment URL (Markdown): [/api/post/concave-utility-question/comments/JpmDQBjqfBZTqMKAN](/api/post/concave-utility-question/comments/JpmDQBjqfBZTqMKAN) * Comment URL (HTML): [/posts/uJnR4YmG5Kq9FfTey/concave-utility-question/comment/JpmDQBjqfBZTqMKAN](/posts/uJnR4YmG5Kq9FfTey/concave-utility-question/comment/JpmDQBjqfBZTqMKAN) Your B3 is equivalent to A4 (assuming A1-3). ### Comment by [Scott Garrabrant](/users/scott-garrabrant) on [Concave Utility Question](/api/post/concave-utility-question) * 2023-04-16 18:45:06Z * Karma: 6 * Total votes: 3 * Comment URL (Markdown): [/api/post/concave-utility-question/comments/HgTBTKoorAdyRCz83](/api/post/concave-utility-question/comments/HgTBTKoorAdyRCz83) * Comment URL (HTML): [/posts/uJnR4YmG5Kq9FfTey/concave-utility-question/comment/HgTBTKoorAdyRCz83](/posts/uJnR4YmG5Kq9FfTey/concave-utility-question/comment/HgTBTKoorAdyRCz83) Your B2 is going to rule out a bunch of concave functions. I was hoping to only use axioms consistent with all (continuous) concave functions. ### Comment by [Scott Garrabrant](/users/scott-garrabrant) on [Concave Utility Question](/api/post/concave-utility-question) * 2023-04-15 08:12:37Z * Karma: 2 * Total votes: 1 * Comment URL (Markdown): [/api/post/concave-utility-question/comments/p8TjxbPZxixz3hfFZ](/api/post/concave-utility-question/comments/p8TjxbPZxixz3hfFZ) * Comment URL (HTML): [/posts/uJnR4YmG5Kq9FfTey/concave-utility-question/comment/p8TjxbPZxixz3hfFZ](/posts/uJnR4YmG5Kq9FfTey/concave-utility-question/comment/p8TjxbPZxixz3hfFZ) I am skeptical that it will be possible to salvage any nice VNM-like theorem here that makes it all the way to concavity. It seems like the jump necessary to fix this counterexample will be hard to express in terms of only a preference relation. ### Comment by [Scott Garrabrant](/users/scott-garrabrant) on [Concave Utility Question](/api/post/concave-utility-question) * 2023-04-15 08:09:20Z * Karma: 3 * Total votes: 2 * Comment URL (Markdown): [/api/post/concave-utility-question/comments/47igAoaWde3Hyfuq7](/api/post/concave-utility-question/comments/47igAoaWde3Hyfuq7) * Comment URL (HTML): [/posts/uJnR4YmG5Kq9FfTey/concave-utility-question/comment/47igAoaWde3Hyfuq7](/posts/uJnR4YmG5Kq9FfTey/concave-utility-question/comment/47igAoaWde3Hyfuq7) The answers to Q3, Q4 and Q6 are all no. I will give a sketchy argument here. Consider the one dimensional case, where the lotteries are represented by real numbers in the interval $\mathcal{L}=[0,1]$, and consider the function $u:\mathcal L\rightarrow [0,1]$ given by $u(x)=\frac{1}{2}-(x-\frac{1}{3})^3(x-\frac{2}{3})$. Let $\succeq$ be the preference order given by $x\succeq y$ if and only if $u(x)\geq u(y)$. $u$ is continuous and quasi-concave, which means $\succeq$ is going to satisfy A1, A2, A3, A4, and B2. Further, since $u$ is monotonically increasing up to the unique argmax, and then monotonically decreasing, $\succeq$ is going to satisfy A5.  $u$ is not concave, but we need to show there is not another concave function giving the same preference relation as $u$. The only way to keep the same preference relation is to compose $u$ with a strictly monotonic function $f$, so $v(x)=f(u(x)$). If $f$ is smooth, we have a problem, since $v^\prime(\frac{1}{3})=f^\prime(u(\frac{1}{3}))u^\prime(\frac{1}{3})=f^\prime(\frac{1}{2})0=0$. However, since, $v^\prime$ must be on some $x>\frac{1}{3}$, but concavity would require $v^\prime$ to be decreasing. In order to remove the inflection point at $x=\frac{1}{3}$, we need to flatten it out with some $f$ that has infinite slope at $\frac{1}{2}$. For example, we could take $f(z)=\sqrt[3]{z-\frac{1}{2}}$. However, any f that removes the inflection point at $x=\frac{1}{3}$, will end up adding an inflection point at $x=\frac{2}{3}$, which will have a infinite negate slope. This newly created inflection point will cause a problem for similar reasons. ### Comment by [Scott Garrabrant](/users/scott-garrabrant) on [Concave Utility Question](/api/post/concave-utility-question) * 2023-04-15 06:58:45Z * Karma: 3 * Total votes: 2 * Comment URL (Markdown): [/api/post/concave-utility-question/comments/uQvtrdhGBTnaFEaSd](/api/post/concave-utility-question/comments/uQvtrdhGBTnaFEaSd) * Comment URL (HTML): [/posts/uJnR4YmG5Kq9FfTey/concave-utility-question/comment/uQvtrdhGBTnaFEaSd](/posts/uJnR4YmG5Kq9FfTey/concave-utility-question/comment/uQvtrdhGBTnaFEaSd) You can also think of A5 in terms of its contrapositive: For all $A,B\in \mathcal{L}$, if $A\succ B$, then for all $0