Do people in the community really not think about bet sizing much? For example, recently heard about https://www.yudkowsky.net/singularity/aibox and the premise just seems a bit weird, a bit of strawman (wow look at the dumb human) and a bit of wrong bet sizing (wow, they are so confident they bet nothing, no downside).
Generally, I would have imagined the starting point of the experiment is you are killed if you let the AI out and if you win you get 10 USD. Then you say “actually we can’t legally do that and probably nobody would take up the challenge so let’s just use capital.” So then you make up an overconfident human who thinks they can’t lose so they bet their entire net worth plus all accessible credit that they can win and if they win they receive 10 USD.
I can hear the objectsion “that is ridiculous! nobody would bet so much for so little upside!”. But that is the entire point of mapping true overconfidence to bet sizing properly: to show how ridiculous it is.
I’m sure there is something interesting that is intended by this thought and actual experiment but to someone coming from markets and ML this just seems like some incentive uncertainty problem not something to do with AI or alignment.
D
It is perhaps useful to first offer my own description of e/acc as an outsider.
Incoherence and opacity and “muddled thinking” is a feature for these kinds of clusters of people. This statement is not a moralization of opaque cluster. Compression is useful for those who find it to be useful!
But it can serve as a Time Jail for certain kinds of nerdy people. It serves as a sales funnel for others. It serves as an Outcome Laundering facility for others (upsides are credited and downsides are washed). This would be the Triad of Downsides that I percieve in the space.
Are there funamental ideas there that are good? Yes, acceleration and regulation trade offs need to be discussed. Any community actively discussing trade offs is automatically a positive contributor to current human population (in my books at least).
But I don’t see a lot of evidence of the broader brand being used to address truly uncomfortable topics and this would be my own measure of true success of a movement. “AI safety” is easy to discuss in comparison to other topics that are just as worrisome.
For example, I recently became interested in LW and found most of the discussions a bit of a labelling/compression game without a lot of truly challenging topics (still very useful of course!). But then I found https://www.lesswrong.com/posts/xqXdDs68zMJ82Dcmt/are-humans-misaligned-with-evolution and this truly impressed me. Both in it’s ability to dance around difficult topics in a kind of subtle vulgate and also interweave ideas with AI and tech problems.
Do e/acc speak openaly about human evolution and the misalignment problem or even holistically about the dual human/ai evolutionary problem? I don’t hear it a lot. For example, if one was a true e/acc believer, one should presumably worry about the worst case where it doesn’t happen before humans lose their ability to create advanced technology. And so you would be looking at human populations and cultures and their various outcomes in terms of both tech (re)production and human (re)production.
However I do not see open and frank discussions of these kind of topics so far. Perhaps I am not searching properly?
I do worry about this, except that my personality is such that I’m frequently taking an opposite position of the agents and finding them ridiculous.
My heuristic is probably something like: if using agents to “think” is taking 5-10x longer than I would to simply write something out, it can’t simply be me aggreeing with whatever they say can it? Am I really deluding myself into thinking I am “shouting at the computer because I agree with it?”
I am completely guilty of outsourcing “clarity and prose” decisions because I largely find that my personal taste is alien compared to what I see coming from other people and I truly would like help in “translating” things into a linguistic form that is the lowest cost for the most people in the target audience. I can well believe that the agents are drifting me into AI slop linguistic space here! But I find I have to constantly argue with the agents over content and ideas and merely retaining my original intent.
High Risk events are typically in the tails of probability and outcome. They are both rare and bad. On top of all the usual problems of deliberating on contentious topics, you ADDITIONALLY have a situation where merely to have productive discussions you likely need a high level of quantitative epistemic “wisdom” but there is not such a high barrier to entry for discussion of these issues (nor should their be, because bad outcomes affect people!).
My heuristic for this has been to travel down two branches: a) focus on singular events that actually happen and study the reaction to that even (i.e. the back propagation to the policy weights in RL speak) b) focus on still bad but less bad and more common events. For example, instead of focussing on terrible crimes, look at fly tipping and antisocial behaviour and non-crime unpleasant interactions.
Bringing this into the AI fold, I was thinking that perhaps instead of hyper-focussing on “AI eradicating humans” or whatever the big X stands for in any discussion, focus instead on things that are already quite bad and are almost certainly happening today with AI and AI platforms.
The example I had in mind here, is that AI and AI platforms appear to be doing exactly what social networks have done over the last decade or two: keep humans on the platform. The h2h problem (human to human) feels like it is massively underserved.
I understand everyone at Big AI X, Y, Z is working furiously to make models better, to stop models from getting dangerer etc … and that nobody is being incentivized to solve the h2h problem. One can come up with lots of explanations for why this is and claim no it is “not the AI” doing this emergently bad thing, but ultimately as AI is used to build it’s own platform and business and accrue more and more data and state I think we should begin to think about this.
It doesn’t sound new and sexy of course because it’s the same bad incentives that dating apps have. But it’s now embedded in a self improving swarm of agents and humans bundled in with capital. So it is kind of a different beast.
This quick take should probably be unbundled into two points. a) the “find a smaller, less contentious, less rare, less bad problem that you can not yet solve” b) AI safety is, at least sometimes, at the wrong level of granularity (the agent and not the institution/legal entity).
I think the context was a bit lost but basically I am new and tried posting something quickly from the phone to see how it works. It was definitely sloppy with typos etc. Pure and raw no AI. Then after rejection I tried clean up with AI and the AI filter. I totally understand that amount of AI slop/spam and the need to have controls.
I think the current workflow for those of us who who AI to iterate on ideas and collect information is to draft with AI and the literally rewrite in our own words in the final form. It is quite bizarre but this actually seems to work.
One of the interesting side effects is that I deep dived into “AI detection” and it is actually quite interesting thinking about what this means for these systems and human thought.
Encounters with Newcomb
A friend recently mentioned (deterministic) Newcomb paradox discussions as we were talking about LessWrong having both recently become aware of the community despite being grad students in computation sciences over 20 years ago and reading these kind of blogs.
I was surprised and in disbelief that a significant subset of the community was engaged in discussing the deterministic case. I turned to various AI to see if they thought this was indeed true and they all seemed to think that yes there was a significant probability of a large proporiton of the community engaging in discussion about this at some point.
My most charitable interpretation is that the verbiage is muddled and there are multiple possible “clear” interpretations of the problem and one might have different probability weights to assign to each clear problem, but within each clear version there should be absolute agreement on what an optimal action is given some objective. This is what we in the industry might call “problem binding” perhaps. And I would expect that the LW community should quickly converge to this kind of multi-interpretation state where there is no real disagreement except on the probabilities for each interpretation.
However, by my quick sweek this does not really seem to be the case. Is it really possible there is a persisten confused/muddled/disagreement state on this problem?
And then I began to think about the construction of the problem and what did an imagined creator intend?
We start with a trivial problem: which do you prefer, 1k or 1M payoff? But then instead of saying that you rewrite the payoff in terms of some complex structure with some stochasticity that is actually “cancelled” out by something in the complexity. You perhaps add some confusing word choices or not. This feels like the classic kind of word problems one might complain about in high school. I understand they serve a purpose (are you able to unpick language) but in terms of content for LW it feels strange unless one jumps to the metaproblem of asking is this actually some kind of filter or experiment on the LW community itself.
I recently read the political reading list of LW and the Eliezer quote of “human evil and muddled thinking intertwine like conjugate strands of DNA” comes to mind. Not that this particular example (if my observations are broadly true) is actually evil, but it does seem to be a bit muddled.
If we resign ourselves to “the purpose of a system is what it does” then are to we to think of this as a kind of time jail, or perhaps just entertainment for certain subset of the LW community? Am I in time jail on the metaproblem?
What am I missing?
I keep trying to post something on Less Wrong and it is either rejected as “too raw / sloppy” or when I work on making it better with AI it is then rejected as “AI generated”.
This seems a bit naive to the possibility of model error or just plain error.
“Make bad things illegal” always sounds good if you remove all the kinds of errors.
Constructively, we do need infra to generally surface “bad outcomes” and engineered opacity complaints as they arise. This would traditionally have been broadcast media and whistleblowers.
Regardless of the “real estate” I would imagine that we should not necessarily expect a zone to be maximally relevant to a person throughout all stages of their journey.
People have forgotten what newspapers were like before personalized feeds.
I am not dismissing the observation merely trying to inject my qualification about the structure and purpose of what I assume are shared zones.
If “front page” is a personal reco algo please correct me. I am new here.
One naive mental model I have with elections like this is that you either have “activists” or “passive growth hackers” at the two extremes.
Activists have values. They speak up to change the language and shape thought. They do not bend to capture what polls say the people like. They do not Gerrymander, they act to make people move!
Growth hackers are the kind of seemingly valueless or compromised type we often complain about. They will say different things to different groups of people, they will be inconsistent globally, they do not expect people to change, instead they chase growth by learning the campaign function.
This is like RL task Vs pure ML where RL task is associated with actions that change world state where Pure ML task is passive in that sense.
I meant to search for if this ontology exists on LW but am on my phone now. Would love to be pointed at a post about this from the past.
D’s Shortform
What biological properties will AI ever have? (1/n) — Death
Biological entities accrue local state and are not copy-pasteable. This is so obvious it goes unnoticed — but it is the entire foundation of what makes death real. You cannot spin up another instance of a person. The state is local, accumulated over a lifetime, and it terminates with the substrate.
Current AI has none of this. Killing a model is a switch. Another instance is identical. Death is not a concept that applies.
This changes when agents accrue local state faster than it can be extracted or cloned — when the IO budget cannot keep up with state accumulation. This already happens. The write-heavy database systems at certain banks became effectively impossible to migrate cleanly — not because the data was gone, but because full extraction became so costly and lossy that it was never actually done. Death of the hardware meant death of the state in any practical sense.
At that point agent death becomes real in the same sense biological death is real. Not a philosophical claim. Just the same underlying structure.
Sexual reproduction is a separate and non-trivial thing — copy-paste plus noise, with all the complexity that entails. That’s next.
I worry the starting mental model is wrong. Ignore AI. Just think of weapons. You have various regimes. They are developing better and better weapons. You partition the regimes into groups where they can communicate and trust each other enough to collaboratively control their collective rate of weapons development. If you are in a trust cluster, you call that cluster your “allies” and every other regime or cluster is effective a potential enemy even if not directly involved in current threats or attacks.
The only solution is to delete all enemies. You can do this by eradicating them or by building trust/collaboration bridges.
The latter is probably quite complicated and is more the realm of markets, Legal Regime formation and monopolies of violence (Empire Building).
I would assume on LW a good subset of people understand the principles of stable regimes. Roughly the idea that given entities who are semi free to choose which regime they would like to inhabit or deal with, you end up choosing regimes along some Pareto curve of fair vs powerful. It should be obvious that you (an entity) prefer regimes that are more fair (to you at least) and are more powerful (to protect you). If a regime is powerful but unfair, another regime can “outbid” them by being more fair and possibly less strong. There might be a lot of interesting cases here but this is merely to show the gist of it.