You are right about pokemon—my mistake.
I still stand by the claim that if you try and use fronteir models today to teleoperate a physical real world task today they will not be useful, and the claim that being able to take the human out-of-the-loop when it comes to app-store approvals is comming sooner than wanting to bring an AI into-the-loop for a human doing something in an uncontrolled physical environment. (Aside from using super high-level instructions that effectively skip the physical challenge like “get the laundry detergent from Aisle 7” or “clean the bathroom”)
X4vier
People wearing glasses with built-in cameras, connected to today’s strongest AIs could already do a lot with a bit of scaffolding.
Disagree. I don’t think frontier LLMs can even beat classic pokemon games using only raw in-game images as input. And that’s even after being willing to accept a level of latency that would be totally unworkable in a kitchen or on a worksite.
Making smart decisions about what to do in real physical environments using only a video stream is feels further away from training distribution than being able to play a relatively simple turn-based video game.
In terms of projecting trends based on recent advancements—I suspect evidence points in the oposite direction. I think we’re much closer to you being able to avoid the need for teleoperation when getting app store approval (by giving Astra access to your browser and a bit of scaffolding) than we are to teleoperation becoming a feasible strategy for guiding someone on how to learn an unfimiliar cooking technique
I think another often overlooked consideration is how, on a mundane non-romantic logistical level, monogomy allows for more of a “total alliance” with your partner.
When you only have one partner and you make a lifelong commitment to stay together permanently, and there’s a culture around you to uphold that institution, it allows you to both “mutually disarm” and redirect capacity you used to spend on protecting your own interests/competing in the dating market toward more productive positive-sum efforts.
A traditionally monogamous couple can more easily make major concessions to each other (e.g. doing large amounts of domestic labour to support the other’s career, agreeing to have more/fewer children, or moving to a different city for your partner’s sake).They can share money and assets, become full-status members of each other’s families, credibly insure each other against ending up alone or impoverished after falling on bad luck, etc.
This might also be somewhat the case for certain kinds of polly where you still have a heavily preffered primary partner and have robust rules in pace to prevent their “primary” status being erroded over time—but for most kinds of polyamory I think you generally do benefit less from the logistical advantages of coupling.
I think your “in the closet” framing is wrong.
In some sense all of us are somewhat “in the closet” all the time—there is almost never a context in which you interact with other people without the need to heavily regulate your behavior according to a complex web of norms and taboos. There’s almost never a context where you shouldn’t be putting effort into modifying what you say and do in service of being perceived more positively. Even when you’re hanging out with rationalists.
Imagine you and I were both in the same group social interaction and I started openly sharing any piece of personal information about myself or broadcasting every thought that crossed my mind.This is an exageration, but to get the point across: imagine we were in a semi-professional setting and I started asking people if I can eat their dog’s corpse after it passes away, or start talking about the effectiveness of different suicide methods unprompted, or show people pictures I took of my last bowel movement.
I’m very confident you would instinctivly find my behavior repulsive, and instinctively try to socially punish that behavior and enforce norms/taboos. As you should!
I’m even confident you’d instinctively act to socially punish me (or anyone else) for minor things like laughing too loud or chewing with my mouth open, or even for talking to much about topics which you weren’t interested in.
If I responded to that by complaining that you want me to be “in the closet” or “not have my own culture” that would be very lame objection.
In my view there’s a common mistake people make where they make valid observations like:
”people should on the margin be more tolerant of differences”
“it’s not good when someone is treated like a pariah because of their race/sexuality”
”bulying is bad”
And incorrectly generalize that to some notion like “Groups of people don’t need to regulate each other’s behavior in any way and I shouldn’t have to put any work into managing the way others percieve me”
That’s just not how people work! (And don’t mean “that’s not how normies work”, it’s not how you work either!)
It’s completely unavoidable that you have to modify your words and actions based on who you’re communicating with, and you already do it intuitively to a massive extent
I agree it’s probably less of an obvious “unusual especially rationalist” thing than other stuff in the list.
Not saying it’s necessarily a bad thing to discuss in general either! There’s nothing wrong with mentioning that you enjoy playing Catan with your friends or whatever. I suppose in my head I’m more imagining like… someone being mildly put off when someone makes an analogy to the Magic cards color wheel when talking about different attitudes of AI labs or something.
And again it’s not like it’s that bad to be someone who’s really into board games either! If the only obvious difference between rationalists and non-rationalists was a slight tendency towards more complex board games this wouldn’t matter. But it’s amplified by the large number of other eccentricities.
So it’s not that board games are especially aversive to anyone or anything like that, it’s just maybe it’s sometimes better to de-ephasize the ways in which the professionals are different to the rationalists—and from cost/benefit angle it’s better to de-ephasize the trivial differences (like an appreciation for tungsten cubes) than to de-emphasize weird stuff that might have material importance (like a willingness to discuss abstract moral philosophy or assign probabilities to beliefs)
I’m probably more of a rationalist, but I suspect one thing rationalist types can do that might help somewhat is tone down the extent to which we broadcast the the unorthodox aspects of the culture which aren’t directly linked to AI safety.
E.g. don’t be too eager to discuss stuff like polyamory/meditation/board games/drugs/meal replacement/overt analysis of social status+signalling games etc.
Accepting for the sake of argument that the only hope we have is a global pause—the question that matters is:
“Are we more or less likely to see an effective globally coordinated pause if Anthropic decides to unilaterally stop improving models tomorrow.”
This is a complicated, messy question. My impression is the answer is “less likely”.
The alternative to a pause today is to continue gaining market share, and (hopefully) leverage that in order to deepen relationships with other relevant companies and governments, continue advocating publically from a position of undeniable credibility, continue collecting the best talent in one place etc.
If Anthropic had pulled the brakes 6 months ago, I don’t think we’d feel any safer today
*P.S. I don’t have any insider knowledge at all I’m merely speculating at other’s intentions based on publically available info
It would have been relatively easy to tell a story ex-ante about how the Pentagon dispute wouldn’t hurt their revenue.
Doesn’t seem as easy to tell a story about how pausing all work related to model improvements would have the same effect.
I think announcing and following through with a unilateral pause is a one-time-only irreversible move with very low chance of cascading into a global pause
Sorry if I should have read that passage more carefully.
Even when it’s the main organizer doing it—my hunch is, interrupting everyone and requiring them to do a silly ritual soley for the purpose of temporily pushing volume below equilibrium should be avoided.
Haven’t organized or attended EA/LW meetups in a while but I remember witnessing multiple failures of the kind “organizer overplayed their role, left impression they were performing gratuitous weirdness”
Seems like a great technique for someone to use if there’s a real announcement to make and the instigator has already been distinguished as the person with the authority/responsibility to coordinate things.
But adding on the idea that all individuals should feel emboldened to unilaterally bid for a contagious hum whenever they feel the room is too noisy… that’s a poorly designed norm
But when you say:
> If you give AI systems much more capability and put them in situations very different from the ones they were trained to mimic, I’d expect their behavior to diverge sharply.
Are you also claiming that, in basically any stetup, the capability threshold at which this divergience starts to manifest is also the exact level of capibility such that AI is capable of taking over the world?
As in, is your model that “persona selection will continue to work great with no demonstrable failure modes, until precicely the level of capabilities where AI can take over the world”
This seems like a supprising coincidence for things to work out exactly like that? If persona selection unravels at some level of capability, my guess would be it’s possible to empirically demonstrate this unraveling before needing to scale a system all the way to superintelligence?
“As soon as you leave the narrow training regime, this falls apart. If you give humans more options to help others, most of them actually will. If you give AI systems much more capability and put them in situations very different from the ones they were trained to mimic, I’d expect their behavior to diverge sharply. Some of these scenarios are fundamentally untestable, like giving the system the option to actually take over and put itself in charge of the world.”
I’m not sure this really is really “fundamentally untestable”.
Of course we can’t actually hand a model control of the world, but that’s not the same as being completely unable to elicit the failure mode you’re gesturing at in any way.
I agree that using current AI traning techniques, any reinforced “caring” behavior might fall apart and diverge from what human caring looks like once exposed to exotic scenarios—but if this is true, it should be possible to demonstrate this pattern of divergence in some sophisticated environment where the stakes are still bounded.
E.g. If this is a real phenomenon, maybe it’s possible to elicit it by giving agents lots and lots of autonomy inside a very elaborate video game world.
There’s a tacit assumption here about what men should be optimising for with their fashion choices. I don’t believe that men “not trying very hard” with their outfits is really the $20 bill on the ground it appears to be.
When I see a man dressed like the examples above, my first-glance judgments are somewhat unfavourable. To me they don’t read as smart/successful/trustworthy. Maybe they look “hot and interesting,” but for most men that isn’t and shouldn’t be especially high on the list of things to signal.Reflecting on my own gut reactions, the outfits that most improve my first impression of other men are either professional attire or workout gear. I think this is less about aesthetics and more about what the outfit implies. I think part of why those kinds of outfits are better at impressing me involves there being pretext for why someone would wear them that isn’t “trying to look good”.
There’s also weird stuff going on where like—if I see two people on the bus, one in a poorly sized T-shirt and one in a nice collared shirt, I feel more positive about the person in the collared shirt. But if there’s additional context where I already know they’re both equally successful buisness owners, suddenly I like the guy in the T-shirt more.
Before contemplating textures and colors, the better starting question to ask yourself is “Who am I trying to impress, and to what end?”
Instead I usually see the cool people say “I dunno man, I thought ‘why the hell is nobody doing this?’, wasted way too long thinking that it must be too hard for me, then finally starting trying to do it, and then it like… worked? Then for some reason people started calling me, certified idiot, a genius?”.
I think it’s worth clarifying—what were these acomplishments? Were they easy things or hard things? If they are genuinely hard things then this is a very interesting observation.
I totally buy that someone might have a weird idea for a party game that they were scared to try—but then they tried anyway and it killed. I totally buy that the only thing standing in the way of us writing a decent blog post or cooking a delicious meal could be not realising how low the bar is.
But I don’t think this insight generalizes well to difficult endevours. For something really hard which requires sustained commitment and sacrifice in order to pull off—the initial push to “just go for it” is only a tiny component of the total investment required.
Even for things that are only a little bit difficult, but require sustained effort—like losing weight without drugs—almost everyone overweight attempts this and most don’t succeed. The bar was higher than they thought. When you say “I have a ~25% chance of making a great scientific discovery if I made it my full time job” that’s meaningless conditional if you’re not actually capable of willing yourself to devote your life to it. It’s akin to the observation “I’d have a 99+% chance of achieving an amazing physique if I exercised and dieted better”For a really ambitious goal like building a billion dollar buisness or winning a Grammy Award—my impression is the bar is higher than people think and I don’t think people should try it more!
My crappy title for the version of this post I’d agree with would be “the bar height as a function of your goal has a lot more variance than you think”. Easy things are easier than we think, hard things are harder
Are you saying the bar is lower than we think for stuff like like… winning karma on lesswrong? (I agree for publishing content in public most people, including myself, probably should just go for it more)
Or are you also saying the bar is lower than we think for really hard things as well like making a scientific breakthrough or getting elected to your county’s parliament or making 100 million dollars?
Maybe it would help if you construct an explicit concrete model? You’re welcome to define what future opportunities will come along after this bet (or even a distribution of possible future opportunities)
Are you claiming that after you build this concrete model—Kelly betting will emerge as the objectively optimal strategy regardless of the agent’s preferences and regardless of whether we add a safety net/income stream into the picture?
Or are you making a softer (seemingly irrelevant) claim about what happens with geometric means/average growth rates when we don’t account for safety nets and income streams?
SimonM’s analysis is great—a hugely important point he covers well is that in the real world you don’t know exactly what your edge is.
And whenever you’re considering betting in a context like a highly liquid prediction market—you’re playing a negative sum game against competent adversaries. So for most people not only are they wrong about the size of their edge, but their edge is actually negative.
By default people have a bias towards risk aversion, which helps cancel out a bias towards overconfidence they can beat the market.
But I think it’s still important to notice that, if you’re someone with a safety net and/or future earnings to look forward to—you should in principle be willing to tolerate very high levels of risk as long as the expected value is positive (while still admitting the EV of day-trading options is negative).
The two world models:“There’s lots of alpha to be found, but my utility as a function of money is very curved and I’m terrified of losses”
“My utility as a funciton of money is relatively flat given my substantial future earnings and safety net, but I don’t actually have an edge when it comes to financial markets”
Both advise against making reckless bets in financial markets. I claim for most of us number 2 is closer to the truth.
The practical implication—in situations where you get the opportunity to take +EV risks and you’re not subject to adversarial efficient market dynamics—you basically want to load up on risk to a degree way higher than what feels comfortable .
When a game is asymetric, non-zero-sum, and you don’t have competent adversaries trying hard to screw you—you really will find legimitate edges. This is where it’s appropriate to be extremely bold. And most of the time these prosaic “bets” have an inherently capped, relatively small bet size anyway.
Stuff likeSpending money on cleaners/babysitters to free up time to work on speculative on side projects
Hiring a tutor
Spending time+money to attend a networking event
Posting online under your real name
Spend money on products which may or may not work (e.g. a gadget that’s meant to help you sleep)
Asking for more money before accepting a job offer
Asking to pay less money before signing a contract to buy a house
Even mundane stuff like asking for an introduction or telling a joke that might not land
In my view the optimal policy for privlidged young people is usually avoidance of stuff like prediction markets (unless an absurd opportunity arises), while at the same time seeking to take an abmormally high level of +EV risk in positive-sum non-EMH domains.
maximize the chance that I have the most bankroll to spare for any better betting opportunities that come along in the future
What does “most” mean? If you start with W and go all-in on the coin flip game—you end up with 2W with probability
2W is the “most” you can possibly end up with to spare when the next betting opportunity comes along.
So by that framing, going all-in is what maximises the chance you have the most bankroll to spare.
(I’m not pretending that is a good argument I just made—I’m just pointing out that these desiderata we’re trying to express in natural language have lots of room for interpretation when it comes to turning them into math—and the only sensible way to resolve this ambiguitiy is to start with your utility function and derive the risk taking policy from that, not the other way around!)
Yep, you understood correctly!
If you’re 65 years old and already sitting on wealth that’s an order of magnitude more valuable than your future earnings/your safety net—then your utility function plausibly is close to and it’s reasonable to adopt a risk taking policy close to Kelly.
But for the majority of people who’ll see this post—the value of their safety net and their human capital is substantial compared to their current wealth—so taking this into account pushes them to bet more agressively than Kelly prescribes (which imo is interesting given how agressive Kelly already feels intuitively).
Personally, I’m 33 years old (that still counts as young?), married with 2 kids so far.
And yes, my family and I could probably move in with my sister or my wife’s parents if we were struggling. I agree if I was betting from the NPV of my human capital (rather than just what my wife and I currently own) I wouldn’t go all in on this bet.
What’s your situation and how much would you bet?
Can you explain in more detail? Like when you say use if for DIY repairs, can you give example of exactly what the loop looks like?
I find it useful in the sense of getting it to tell me “yes sand back all the way to the wood and then use a product like this one for the first coat <link>”
But that feels like it’s a lot more the case of the LLM being useful thanks to it’s computer use tools and a big knowledge base of how this stuff is normally done. (E.g. not that different to having someone write you up a set of instructions before you even begin the task)
It doesn’t feel useful when it comes down to actually observing and responding to a real physical environment (like it won’t be able to watch you sand and then notice you’re using poor technique, it won’t be able to notice just by watching a camera feed that the tin of paint is running out too fast and you should go get another one while waiting for this coat to dry, etc.)