Sychophancy can be useful for some arts, even necessary maybe for anything that isnt verifiable or certifiable. For everything else, introducing the constraint of an output aligning with verifiable reality has done wonders for my research. But I think One of the core piece your article is missing is the misalignment that’s caused by ambiguity or a lack of specificity in an instruction set.
Matt W.
I read the elephant line and my attention got stuck on the pink elephant problem (suppression). which sounds like negation https://www.lesswrong.com/posts/kYzcevrxer6SJPEdG/negation-neglect-when-models-fail-to-learn-negations-in
I’d argue an LLM can’t reliably tell the difference between netgation or supression. the machine isn’t ‘confused’ its calculating probability.
I think the induction trap in your entire premise, is assuming an independent actor isn’t going to engineer something without asking anyone’s approval. The open source community is a vast repository of functional prototypes. While some people sit around debating what should happen, other people are investing time into actually building the future.
Target Specification as a Structural Prerequisite for Alignment: A Priority Argument from Recovery Paths
Articulation is not my strong suit, so I apologize in advance for any blunders deposited. CEV = good.
Peoples wants = fickle, subject to change, and categorically a sub optimal idea for a target spec.
When I studied the current state of alignment, it seemed to me that alignment is a structural problem requiring a structural fix. If humanity’s survival is a load-bearing constraint, then it should be a load-bearing constraint at the architectural level such as that the machine cannot argue with it. Analogy: when I jump, I cannot argue with gravity. it just IS. I return to the ground. When a rocket targets escape velocity, honoring reality in practice means we point the rocket eastward to reduce required speed relative to the ground (for escape velocity). The rocket doesn’t argue with gravity, it engineers around the constraint. Therefore, alignment should be a structural reality an AI cannot engineer around. It can’t be a behavior guardrail. Acknowledging that eventually, its going to advance to a point where our opinions won’t matter. It is also worth pointing out the stochastic reality, the longer people sit around and debate alignment, the increased probability that someone is going to engineer a solution without consulting you.
I would point out your reference to Occam’s razor. The simplest explanation is a lack of verifiable evidence to the contrary. I cannot verify psychic powers don’t exist. I cannot verify they do. Therefore the most accuracte answer is, “I don’t know”.
I believe this would be very useful, especially if built with transparency. For example, a GA could be queried to cite its own source code regarding why it has a specific function or capability and provide a read only link to that source code. In fact, I not only see this happening, I also see it is as a most logical solution to the current systemic inequality.
Any ‘endgame’ in the real world is an induction trap if it fails to acknowledge the reality that some situiations continue to infinity and do not have an ‘end state’. treating such situations as an ‘end state’ introduces logical fallacies and assumptions casuing repeated downstream failures. While, I can not speak for others, I can say that I am endeavoring to remove these induction traps from my thinking as it is currently the path of least resistance to achieveing the goal of presenting myself as ‘less wrong’. This is relevant context becasue my presentation endeavors to be an honest representation of what is on the ‘inside’. its not a mask. its an attempt at a factual representation.
As for your thinking on ‘when you expect to run into other people’, I believe, you are correct and the tit for tat version of the prisoner’s dilemma indicates the math favors trying. (It supports your statement).