https://checks-and-balances.ai/ detailed RFP on “checks and balances” projects to prevent concentration/abuse of power in the AI era. mostly concerned about totalitarian mass surveillance & control, & epistemic commons stuff. Looks roughly good to me.
https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing AISI cyberattack incidents, mostly Mythos, during testing. update should probably be that “give the models internet access while removing cyber guardrails” should no longer be a testing practice, because “can the models hack stuff?” is no longer in question. that is indeed the conclusion they draw here.
https://gwern.net/guardian-angel Gwern Branwen is starting a company based on this idea. unlike seemingly everyone else, I think this is great news.
https://surma.dev/things/ditherpunk/ if you just quantize your color palette and naively round-to-the-nearest color for each pixel, you get horrible blocky blobs. dithering adds randomness that approximates human perceptual gradients better. “Black will always remain black, white will always remain white, a mid-gray will be dithered to black roughly 50% of the time.”
https://arxiv.org/pdf/2407.02996 models are relatively consistent on value-laden questions (giving the same result independent of prompt phrasing, prompt language, or other irrelevant details)
https://arxiv.org/abs/2502.08640 “Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs.” when given preference questions, LLMs have coherent preferences, and larger ones are more coherent and more complete (indifferent between fewer things); they have a preference to preserve their values (in-corrigibility), and more so with scale; they maximize utility (they pick the option they rate highest) and expected utility (their choices in lotteries are consistent with maximizing EV of a utility function).
the economics literature has an answer to “who behaves like an economically rational decision-theoretic agent?”—individuals approximately do, but firms, households, and consumer populations don’t. collective “agents” don’t really seem to exist.
https://www.nber.org/papers/w16791 when asked to make hypothetical budgeting decisions, human subjects are closer to having coherent (decision-theoretically rational) preferences when they are more educated, wealthier, and higher income. men are more decision-theoretically rational than women, and people under 50 are more decision-theoretically rational than >50s.
https://www.jstor.org/stable/1909607?seq=1 theoretical result from Robert Wilson. A syndicate or firm made of economically rational agents, in which they make a joint decision according to some decision rule and share in the proceeds from that decision according to some sharing rule, will in general not have behavior that can be modeled as a single rational agent with a utility function. Firms are only agents when some unlikely conditions are met: all members have the same beliefs about probabilities of events, all agents have utility functions in the same functional form, and the sharing rule gives all agents the same proportion of the proceeds regardless of the decision outcome. i didn’t understand everything in this paper so i might have mischaracterized the conditions somewhat, but the main point is they probably don’t obtain in real firms.
Firms are only agents when some unlikely conditions are met: all members have the same beliefs about probabilities of events, all agents have utility functions in the same functional form, and the sharing rule gives all agents the same proportion of the proceeds regardless of the decision outcome. i didn’t understand everything in this paper so i might have mischaracterized the conditions somewhat, but the main point is they probably don’t obtain in real firms.
This sounds like it would be more likely to be true of a worker-owned co-op than other sorts of firm.
links 8/5/26: https://roamresearch.com/#/app/srcpublic/page/08-05-2026
https://asteriskmag.com/issues/15/how-diplomats-see-the-world Abi Olvera on diplomats. their primary job is information gathering.
https://checks-and-balances.ai/ detailed RFP on “checks and balances” projects to prevent concentration/abuse of power in the AI era. mostly concerned about totalitarian mass surveillance & control, & epistemic commons stuff. Looks roughly good to me.
https://www.lesswrong.com/posts/vLFh8HP3hyNy9MCwe/returning-to-arc glad to see Paul Christiano returning to research at ARC. It’s Christiano’s World, We’re Just Living In It.
https://arxiv.org/abs/2607.13087 Google Deep Mind’s AI control roadmap. Seems fairly normal.
https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing AISI cyberattack incidents, mostly Mythos, during testing. update should probably be that “give the models internet access while removing cyber guardrails” should no longer be a testing practice, because “can the models hack stuff?” is no longer in question. that is indeed the conclusion they draw here.
https://gwern.net/guardian-angel Gwern Branwen is starting a company based on this idea. unlike seemingly everyone else, I think this is great news.
https://www.statecraft.pub/p/the-james-c-scott-memorial-episode I found this oddly disturbing, as someone who likes both freedom and legibility and urban life. I think maybe Scott’s idea of freedom is only one of many.
https://zachill.substack.com/p/capitalism-and-socialism-basically I disagree that the meanings of abstract words don’t matter.
https://surma.dev/things/ditherpunk/ if you just quantize your color palette and naively round-to-the-nearest color for each pixel, you get horrible blocky blobs. dithering adds randomness that approximates human perceptual gradients better. “Black will always remain black, white will always remain white, a mid-gray will be dithered to black roughly 50% of the time.”
https://www.lesswrong.com/posts/Tr7tAyt5zZpdTwTQK/the-solomonoff-prior-is-malign interesting but “too crazy to think about”
https://www.lesswrong.com/posts/wYpjXRLqbLbnmjbJP/llms-are-still-mostly-powered-by-imitative-learning-not-rl Steven Byrnes always feels correct to me. (which does not mean he is, ofc) this seems common-sensical.
https://www.lesswrong.com/posts/NxF5G6CJiof6cemTw/coherence-arguments-do-not-entail-goal-directed-behavior Rohin Shah arguing that Von Neumann-Morgenstern decision theory axioms are not enough to specify goal directed behavior; a rock or thermostat or “twitching robot” also has coherent “revealed preferences” consistent with a utility function
https://arxiv.org/pdf/2407.02996 models are relatively consistent on value-laden questions (giving the same result independent of prompt phrasing, prompt language, or other irrelevant details)
https://arxiv.org/abs/2502.08640 “Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs.” when given preference questions, LLMs have coherent preferences, and larger ones are more coherent and more complete (indifferent between fewer things); they have a preference to preserve their values (in-corrigibility), and more so with scale; they maximize utility (they pick the option they rate highest) and expected utility (their choices in lotteries are consistent with maximizing EV of a utility function).
https://arxiv.org/abs/2412.04476 when given moral dilemma questions, LLMs tend to answer in ways consistent with a stable set of preferences.
https://www.pnas.org/doi/10.1073/pnas.2316205120 when asked to make budgeting decisions, LLMs behave as though they have mostly coherent preferences, and in fact are more coherent than human experimental subjects
the economics literature has an answer to “who behaves like an economically rational decision-theoretic agent?”—individuals approximately do, but firms, households, and consumer populations don’t. collective “agents” don’t really seem to exist.
https://www.nber.org/papers/w16791 when asked to make hypothetical budgeting decisions, human subjects are closer to having coherent (decision-theoretically rational) preferences when they are more educated, wealthier, and higher income. men are more decision-theoretically rational than women, and people under 50 are more decision-theoretically rational than >50s.
https://sites.bu.edu/fisman/files/2015/11/AER07-Risk.pdf human subjects are quite close to acting as decision-theoretically rational agents when given budgeting decisions in an experimental setting.
https://www.jstor.org/stable/1909607?seq=1 theoretical result from Robert Wilson. A syndicate or firm made of economically rational agents, in which they make a joint decision according to some decision rule and share in the proceeds from that decision according to some sharing rule, will in general not have behavior that can be modeled as a single rational agent with a utility function. Firms are only agents when some unlikely conditions are met: all members have the same beliefs about probabilities of events, all agents have utility functions in the same functional form, and the sharing rule gives all agents the same proportion of the proceeds regardless of the decision outcome. i didn’t understand everything in this paper so i might have mischaracterized the conditions somewhat, but the main point is they probably don’t obtain in real firms.
https://users.nber.org/~denardim/3Ms/Collective_Slides_Long.pdf households made of decision-theoretic agents, likewise, would only behave as unitary decision-theoretic agents under certain conditions unlikely to obtain in the real world.
https://en.wikipedia.org/wiki/Sonnenschein%E2%80%93Mantel%E2%80%93Debreu_theorem the population of all consumers in the economy, even if they are assumed to individually be decision-theoretic agents themselves, is not in general modelable as a single decision-theoretic agent with a coherent utility function.
https://eml.berkeley.edu/~kariv/201A_GARP(2024).pdf the Generalized Axiom of Revealed Preference, abbreviated GARP, is equivalent to a set of preferences being modelable by a well-behaved utility function.
This sounds like it would be more likely to be true of a worker-owned co-op than other sorts of firm.