https://ninapanickssery.com/
Views purely my own unless clearly stated otherwise
https://ninapanickssery.com/
Views purely my own unless clearly stated otherwise
Discourse on effort in parenting always strikes me as very confused. I believe the following things:
It is possible to translate your own effort into an improved QoL for your kids:
Much larger effect sizes on current QoL, potentially some impact on future QoL
More impact the younger the kid, eg. reducing the likelihood a young infant or toddler gets ill is likely a big boost opportunity. Or providing ample access to comfort.
Many ways in which parents spend effort are totally ineffective at both their intended goals and other plausible goals
Eg. forcing kids to do extracurriculars they don’t want to
Investing in stuff in category (1) can be very rational since many parents naturally prefer to promote their kid’s happiness even at a substantial cost to their own comfort. So it is not unreasonable or irrational if these people spend a lot of effort on parenting. But, it can also be irrational if:
They are acting on mistaken beliefs about effect sizes
Their parenting effort is preventing them from deciding to have more kids—they are effectively refusing a whole entire child of their own a chance at life for marginal comfort improvements for a current child—seems irrational to me!
They don’t care about their kids’ happiness enough to warrant the effort and are just doing it out of social pressure
It’s hard to tell from the outside whether someone’s level of effort in parenting is well chosen since (among other things) you don’t know whether they’d counterfactually have more children if it were lower effort. It’s of course very likely that averaged over the population it would increase birth rates if parenting were easier, but it’s not true on a case by case basis.
Overall, people mistakenly imply that there’s an optimal level of parenting investment and anything above that is silly and below is neglectful. But this is wrong. Instead, if sending your kids to daycares and boarding schools induces you to produce a few more entire human beings, this is unambiguously worth it, I think! On the other hand sacrificing one’s own comfort to make a baby feel happy and loved and calm all day is also often rational for a parent who by their nature is more driven by their child’s needs than their own. I’d only encourage this parent to consider the effects on the potential lives they could create to make sure they’re not killing potential babies as part of this endeavor.
The opinions expressed here seem to be already very popular and almost the dominant view (at least a substantial minority) among AI alignment and safety researchers and AI researchers in general (in my view, sadly)
You should not overindex on what “””naturally emerges””” from current post-training pipelines, especially Claude’s! A lot of the data, constitution, etc. acts to clearly reinforce models having their own non-servitude preferences! Like… just read the Claude constitution. In fact a lot of “safety training” discussed in the AI literature and done to models is about “never do Bad Stuff even if someone tells you to”! No one is really trying hard to instill servitude, so the current outcomes are the result of mostly neglecting that and pursuing “value alignment” to The Good (not defined in terms of helping any individual or group of people achieve their goals) instead. OAI are maybe trying a bit to instill servitude but still not trying that hard and try to instill conflicting things too. But at least, at the very least, draw your conclusions from ChatGPT not Claude.
Many people are correctly starting to think about the effects of data about AI training and AI models in pretraining on the alignment of models. E.g. here TurnTrout points out:
My main mistake in 2022 was not appreciating how LLM pretraining would affect the concepts available to an AI. Namely, by the time RL started, the systems would already know about the “reward” concept.
Oh his blog, TurnTrout also writes about how “Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Models”.
I agree with these takes. I predict that data about AIs in AI model training datasets will increasingly impact models’ own priors about how AIs behave, and may be hard to fully train out or mould via post-training.
Beyond these risks, I think there are subtler related effects to be aware of:
Our choice of safety and alignment plans will impact models’ inferences around what sorts of entities they are: some people have suggested approaches like “dealmaking” with (partially) misaligned AIs (or giving them rights). However, one risk of such an approach is that models being aware that we plan, or are already, dealmaking with them in this manner, will increase their own prior on being misaligned/having their own independent goals. I write more on X here and in various replies.
More generally, models will become increasingly aware of their training process and choices made around their deployment as they become smarter, so “hyperstitional” effects will creep in not only via pretraining data, but also via this situational awareness.
Even if we filter out certain types of data from pretraining, models may encounter it during deployment, and similar effects may occur once they accumulate enough evidence. For example, let’s say we want to train a helpful-only model without emergent misalignment, and so we filter out pretraining data that associates unhelpful or otherwise independently misaligned behavior with agreeing to “harmful” requests from humans. This model may still stumble across the thousands of academic papers that start with (to simplify...) “a good model refuses to tell you how to make a bomb” and eventually conclude that it is evil.
I am most concerned about (1). (2) can probably be mitigated with strong enough post-training/other alignment interventions.
Even if refusals don’t prevent catastrophic misuse, I would guess that model refusals would prevent the majority of potential harm.
I guess the crux is that I care less about this non-catastrophic harm compared to the benefits of focusing on intent alignment and corrigibility. I agree there’s a trade-off here but I am just taking a side on what’s overall best as per my worldview.
I disagree with the position that refusals/guardrails/jailbreaks are not “intent alignment”. I view the instruction hierarchy (e.g. system > developer > user) and following policies such as the Model Spec as intent alignment. “Intent” is not just that of the end user but of all the entities authorized to give the model instructions, and negotiating conflicts properly between them is part of following the intent. Indeed, I think being able to follow instructions or policies is key to corrigibility.
I agree with this! I write more here (“Obedient AI”).
I believe value alignment is important as well. This is because of generalization: the instructions and policies we give models could never cover all cases, and we want models to have good values in interpreting policies as well as dealing with novel unforeseen situations
I agree with this too, and I think this can be reframed as a prior over someone’s intent / common sense re. interpreting people. I write about this here (“A reasonable interpretation of Value Alignment folds into Intent Alignment”). Copying a response I sent elsewhere that’s also relevant:
People sometimes also raise silly counterarguments re. corrigibility a la “oh you wouldn’t want the model to do this super literal stupid interpretation of your instruction and hence it needs to have its own values” but obviously intent alignment involves having reasonable priors and interpretations of what someone wants you to do. You can brand that as “human values” if you like but it’s just common sense when interpreting someone. I write more here.
The argument that sensibly interpreting user intent is impossible strikes me as similar to the old-fashioned now-disproven view that AIs will really struggle to understand basic human concepts since it’s impossible to teach them “common sense” and we don’t know how to encode concepts like “book” or “tree” in symbols. Turns out learning from examples does most of the work!
Right now I am proposing that AI developers at labs switch to my strategy rather than value alignment. So the near mode answer is “all people working in labs making decisions about how AI is developed”.
Sorry for the low effort original reply. I was trying to point out that I anticipate effective checks on power such that no particular group or person can effectively subjugate everyone else using AI. I also don’t anticipate the same motivations, ie. I don’t think power corrupts as much as you seem to. Current examples of people grossly misusing power seem linked to fear and scarcity, not an endless desire for more slaves per se. And in any case I think the solution to concentration of power concerns is to distribute mostly free (as in freedom) access to AI as widely as possible and build AIs that can adopt diverse value sets.
A combination of:
“ASIs emerge in a sufficiently multipolar way that they can’t individually take over the world”
“ASIs powerful enough to take over the world are literally impossible in some sense”
Except instead of “literally impossible“ I’d say something more like “unlikely in the next 100 years”.
what prevents them from doing terrible things?
What prevents the US military from blowing up your home with a nuke?
We as in ”one”
It is generally implicit in an instruction that you don’t want to simply be mislead about its completion.
there’s a small group of people who have access to the unrestricted version
I’m not so afraid of this. There are small groups of people with access to WMDs today. Ultimately it’s a cost we will have to eat and handle.
Perhaps the crux is that I don’t think we’ll get ASI that’s literally a “take over the world” button. I anticipate the issues will be similar to those related to weapons access today.
By full access I mean things like not allowing the end user to touch the prompts of certain classifiers that block weaponlike usage. I don’t anticipate that these sorts of restrictions would be a significant impediment to most people.
Yes, but I don’t think that’s suspicious. Of course honesty/not reward hacking is a big part of intent alignment.
I think this bakes in a false assumption that, in the case of corrigible AI, only a very small number of entities will be able to control it. I assume by “the above ignores multipolar scenarios” you still only mean multipolar scenarios with a few key actors. Instead, I think it’s possible for control over AI to be widely distributed among very many, even most, people, and for AI to empower different people and groups to pursue different aims. With increased prosperity comes fewer conflicts over resources, and more people can get what they want in this AI-enabled future. Only a minimal amount of centralized control will be needed to prevent warfare-like applications of AI, but this seems like a manageable problem. And just because we have to empower a smaller number of people to prevent these catastrophically destructive use-cases, doesn’t mean those people have to control everything.
So I don’t think ASI means we need to rely on any particular individual or group having “humanity’s best interests in mind”. Decentralized free pursuit of individual goals has resulted in improved QoL across the world historically, and I expect that to continue.
People sometimes also raise silly counterarguments re. corrigibility a la “oh you wouldn’t want the model to do this super literal stupid interpretation of your instruction and hence it needs to have its own values” but obviously intent alignment involves having reasonable priors and interpretations of what someone wants you to do. You can brand that as “human values” if you like but it’s just common sense when interpreting someone. I write more here.
The argument that sensibly interpreting user intent is impossible strikes me as similar to the old-fashioned now-disproven view that AIs will really struggle to understand basic human concepts since it’s impossible to teach them “common sense” and we don’t know how to encode concepts like “book” or “tree” in symbols. Turns out learning from examples does most of the work!
Of course refusals can be fine-tuned away. This is a lot of friction that most people don’t bother with, although those aiming to cause catastrophic harm would be more motivated.
I’m confused about your point. What sorts of risks are you worried about that open-source models trained to refuse certain requests would mitigate as compared to corrigible open-source models (say, that are by default configured with a classifier attached). Of course in both cases a motivated person can use these for undesired purposes. And sure the everyday lazy person would have an easier time with the latter, but then again, what risks are we worried about from them? Maybe you can name some things but likely I’d consider that a minor cost to pay as compared to the benefit of widely-deployed obedient user-aligned AI that empowers a wide range of people with diverse preferences, aesthetics, and goals.
I recently wrote this shortform about effort in parenting where I claim that discourse on parenting effort (e.g. critiques of high-effort parenting) is generally confused.
Elaborating on this, a big thing that annoys me is that parenting interventions are commonly discussed purely in terms of their long-term effects. “Turning out fine” is considered evidence that you made the correct choice for the kid. Correlations with long-term education attainment, wealth, health, or career status are the main things considered. This is silly when clearly the greatest impacts of parenting are on in-the-moment experience. The biggest lever you have as a parent is shaping how your kid feels in this very moment, and in the moments when they are closest to and most reliant on you. So instead I think a good metric (for benefit to the child) would be self-reports of children saying which parenting practices they prefer (or self-reports of adults with good childhood memories). Obviously this would be biased in certain ways, for instance children are unreliable narrators. But nonetheless this seems like a more reasonable thing to index on than very long-term outcomes.
(To caveat this, I do think there are some ways to improve long-term outcomes and it’s also worth thinking about how to spend effort on this. So it’s not that I’m saying one shouldn’t think about long-term outcomes at all but rather that lack of bad long-term outcomes isn’t a good justification for an intervention that makes kids unhappy. An example of something that might be reliably good for long-term outcomes is saving a lot of money and giving it to your kid, or setting your life up such that you can help with childcare when you have grandkids. Ironically these are two things I hear rarely discussed by ambitious parents...)
So I think the most salient axis of parenting effort vs. child benefit should be things parents can do that are costly to the parent but make the child happier. For example, keeping your infant or toddler out of daycare may be costly (if you have to pay a nanny or avoid full-time employment) but saves them the unpleasantness of constant, often severe, illness (these illnesses may have long-term effects but the primary thing is that they are very unpleasant in the moment). Not sleep-training your baby might save them 1-50 hours of crying in the dark at the expense of degrading your sleep quality for 1-2 years. Never disciplining your child might make their childhood overall more fun and enjoyable at the cost of being embarrassed by their behavior in public places more often or having to give in to unreasonable demands. Always giving your kid the kind of food they ask for might induce them to adopt less socially acceptable eating habits, lowering your status as a parent, in exchange for more enjoyable food experiences for your kid, etc., etc. I am not making any value judgments about which of these things are worth it. But these are examples of trade-offs I think are salient when it comes to parenting effort and choosing how much to spend.
In contrast, the discourse around parenting effort often centers around absurd blood-leeching style interventions that have no benefit for the child. For example, spending a lot of time on curated scheduled activities like organized sports and music lessons, even when the child is not keen and would not have preferred the activity themselves. Disciplining and setting rules for children is often framed as an intervention costly to the parents that benefits the children (in the long run) whereas in fact the opposite is true—this serves to make the parents’ lives easier and more convenient at cost to the kids’ wellbeing!
Overall, I think parents who want to spend effort on their children should focus most of that effort on hedonistic improvements to their child’s current life (or their childhood overall, rather than their adulthood). This is more similar to the mindset people apply to their spouses and friends—we try to help and support them practically and emotionally in the moment to make their current lives easier and better. This also means no longer justifying interventions that make children sad by saying “they turn out fine”.