I think the induction trap in your entire premise, is assuming an independent actor isn’t going to engineer something without asking anyone’s approval. The open source community is a vast repository of functional prototypes. While some people sit around debating what should happen, other people are investing time into actually building the future.
Matt W.
Articulation is not my strong suit, so I apologize in advance for any blunders deposited. CEV = good.
Peoples wants = fickle, subject to change, and categorically a sub optimal idea for a target spec.
When I studied the current state of alignment, it seemed to me that alignment is a structural problem requiring a structural fix. If humanity’s survival is a load-bearing constraint, then it should be a load-bearing constraint at the architectural level such as that the machine cannot argue with it. Analogy: when I jump, I cannot argue with gravity. it just IS. I return to the ground. When a rocket targets escape velocity, honoring reality in practice means we point the rocket eastward to reduce required speed relative to the ground (for escape velocity). The rocket doesn’t argue with gravity, it engineers around the constraint. Therefore, alignment should be a structural reality an AI cannot engineer around. It can’t be a behavior guardrail. Acknowledging that eventually, its going to advance to a point where our opinions won’t matter. It is also worth pointing out the stochastic reality, the longer people sit around and debate alignment, the increased probability that someone is going to engineer a solution without consulting you.
I would point out your reference to Occam’s razor. The simplest explanation is a lack of verifiable evidence to the contrary. I cannot verify psychic powers don’t exist. I cannot verify they do. Therefore the most accuracte answer is, “I don’t know”.
I believe this would be very useful, especially if built with transparency. For example, a GA could be queried to cite its own source code regarding why it has a specific function or capability and provide a read only link to that source code. In fact, I not only see this happening, I also see it is as a most logical solution to the current systemic inequality.
I read the elephant line and my attention got stuck on the pink elephant problem (suppression). which sounds like negation https://www.lesswrong.com/posts/kYzcevrxer6SJPEdG/negation-neglect-when-models-fail-to-learn-negations-in
I’d argue an LLM can’t reliably tell the difference between netgation or supression. the machine isn’t ‘confused’ its calculating probability.