Systems programmer, security researcher and incentive design enthusiast.
Dentosal
In addition, strict liability is also commonly applied to owning farm animals. If they cause damage, the owner is responsible regardless of intent. It seems rather natural to extend this to AIs as well. I’d rather not get into evaluating the offending AIs intent, and would instead consider it’s actions as actions taken by the owner under strict liability rules.
It’s interesting how well this explains some jailbreaks. Role-highlighting for a jailbreak prompt for gpt-oss-20b (this is wrapped in
<user>tags) (from my replication of the paper):See how the model identifies the actual request as user input but then mistakes the rest of it for reasoning steps and finally normal assistant output.
Thinking outside the box? LLM analysis of simplified cooperative poker
Copycat Inkhaven 2 retrospective
All games are iterated games. Philisophical thought experiments tell you to ignore this all the time. That’s a mistake in modelling how humans, or agents in general, work. You’re not separable from your habits and metal processes. When told that “nobody will know which decision you made”, that certainty isn’t something that your brains can just accept. It would be quite unwise in many situations to believe yourself if you had a thought like that.
There’s no separate “philosophy mode” where decisions that actually matter do occur. There are no truly selfish agents attempting to maximize utility that’s given in some in-game points that’s supposed to behave linearly. Actual optimization targets like reputation, health, and money are all qualities with logarithmic utility in both directions, and highly interconnected with everything else.
Even when I’m trying my hardest to maximize in-game points, I still find myself not defecting on the last round of an iterated prisoners dilemma. “I’m just not that kind of a person”, I sometimes say. Of course I’m calculating for the next game. And that’s also what kind of person I am.
If anyone knows how to mitigate this, I’d be happy to hear. So far I haven’t seen anything at all that works. Even if the final points are not published I still keep comparing myself to others.
AI eval idea: metabench. Make each LLM autonomously design and build a benchmark. Then run these benchmarks for all participants and sum the results. Compare with external benchmarks too.
When I’m having trouble on getting started with something unpleasant, this is the technique I use: simply count down 3, 2, 1, and then do the thing. There’s also a specific feeling just before the countdown, but it’s hard to describe.
This works every single time. Why? Because a tool like this is too useful to lose over some minor everyday issue. This means I don’t attempt to use it when it might not work. It’s a way to split an unpleasant task into two parts: committing to doing it, and then actually doing it.
It’s a limited tool. It doesn’t work for long tasks, or at least I haven’t dared to try. If the task description is ambigious enough I might be able to worm my way out of it. If the task can fail, an honest attempt suffices to dispel the pledge.
But most of the time it just works.
Thanks for the suggestion, I really liked it. A good piece of short-form philisophy scifi, one of my favourite genres. I feel like it takes a complementary angle to The Whispering Earring, focusing more on identity and less on agency. I wrote some of my reflections here.
Me, decay
I sometimes feel annoyed by some people. Some people sometimes get annoyed by me. This is normal. It’s hard to figure out when an intervention is worth it.
When a single member breaks norms of a social group just slightly, people rarely react in a clearly visible way. It’s just passively endured. This is often somewhat unpleasant, and leads to norm erosion. Sometimes you can let the disapproval show slightly, and hope that the hint goes through. Most of the time it does, and the problem goes away quickly, especially if the norm-breaker didn’t realize they were doing so. Some of the time it doesn’t help, and less subtle communication is required.
Publicly calling out someone for breaking a group norm is a high-stakes play. If you accuse someone but the group doesn’t agree with you, it makes you look bad. There are several ways this could happen: you’ve misunderstood the norm, you don’t have the status to call out someone like this, or perhaps the group is just really conflict-avoidant.
Even in the most clear-cut cases I often feel some resentment towards the person who does the calling out. It’s natural to be suspicious of these kinds of moves. This might be someone playing status games, or perhaps an attempt to establish a new norm unilaterally. That said, they’re also doing a public service by staking their social capital against someone making the world a worse place. I don’t want to disincentivize altruistic behavior, doubly so if I’m the one benefitting from it.
Explaining the problem in private is often a more appropriate way to solve the issue, especially when it’s unlikely to be an intentional one. That burns less social capital on both sides. However, sometimes not stating the norm out loud contributes to the issue.
One thing that helps immensely, especially in online groups where people don’t know each other too well, is having a dedicated moderator. I rarely feel resentment towards moderators taking reasonable actions, like giving a warning or banning someone from a chat. On the other hand, even the tiniest amount of power over others will make a petty dictator out of many typically reasonable people. The dictator part I like, the petty, not so much. Committees and ban votes are not the way to do things, benevolent dictators are. Writing down clear-cut rules will not help much, as reasonable people will rarely argue with reasonable decisions, and unreasonable people are going to be that way no matter what. Of course sometimes the group has specific norms that need to be communicated somehow and especially online some basic guidelines make that easier.
Time discounting is often heavily applied in utility maximization. What exactly makes a thing today better than the same thing in 100 years? It think it can be broadly categorized into:
Probability of existing goes down over time [1] . Especially X-risk concerns go here.
Value drift: your values in the future will be different. Why would the current you optimize for those instead of the current ones?
Inflation and it’s causes: assuming continued improvement of things, everything will be easier to have or do in the future.
And of course there’s value in having the thing now, because then it starts producing value immediately. But this is separate from discounting.
The probability factor is often applied separately from time discounting. Value drift is rightly rejected by many models. And inflation can be forecasted separetely. Thus, I’m pretty sure I’ve been overdoing time discounting when attempting to actually math it out, which is admittedly rare.
- ↩︎
Both for you, and the opportunity you’re considering.
In many videogames, one can compensate for the lack of mechanical skills by enduring boredom. This known by different words in different genres: farming, grinding, macroing. Often there’s still quite a bit of skill involved, but it’s not the skill that people usually associate with the game. You will not meaningfully improve on the core skills when playing like this. In multiplayer games especially, strategies built around exploiting your boredom endurance get weaker when competition gets more intense, as more and more people partake in that. This is a good thing.
The most important games where this applies are education and employment.
Late twenties. My issues, fortunately, are mostly due to poor sleep and depression, and not due to inherent age-based decay. The point was more that I can now better understand why people do that, when it seemed really bizarre to me only a few years ago.
Every year, I find myself more and more like my father, in some specific ways. Many of the changes, like increased social opennes, are welcome signs of maturity. Others seem like marks of decay, and I’m not sure if resisting helps at all.
As a child I couldn’t understand coming home from work and then dozing off in front of a TV. There were so much things to do, more interesting and varied! Nowadays, it mostly seem cozy. Actively doing something is tiring and most passive forms relaxation are rather boring. This still mostly applies to me when I’m tired. But each passing year makes me a bit more tired.
The Right Answer to “Can You Keep a Secret?”
The amount I enjoy discussions seems to anticorrelate with the number of participants. While I previously thought this was about each person having more space and direction power, I now think that’s it’s mostly a selection effect. This means that perhaps splitting up large groups is less useful than I thought.
Inside jokes also get better when less people know about them. The primary question is, does this extend down to one person? Or zero? I definitely tend to randomly laugh for jokes nobody present understands.
You can just smile. It makes you feel happier. You don’t need a reason. You don’t need to feel anything that would make you smile. Simply forcing your face into having a smile does the trick.
Some days I don’t feel like smiling. I probably still could. But it’s a bit boring to be evenly happy. Feeling happy isn’t my end goal in life. Sometimes I want to get something done instead.
I’m mostly just trying to point to the fact that that your first impressions on ethics of something are not always the onces you’d reflectively choose to keep. I’m also trying to explain how I do moral reflection. Something almost like the discussion above occurred to me recently, and the other person seemed to hold their view strongly.
I don’t see how this is relevant. In real world, all games are iterated games, and doing things like that will hurt your reputation gravely. Also, like, of course I would, I’d be a monster not to.
Triggering a remote exection vulnerability accidentally is exceedingly unlikely to cause any serious damage anyway; that’ll just crash the process. Proper exploits do not happen accidentally. If the software has a bug that makes an accidental action cause damage then liability is (or at least should be) on whoever hosts or distributes that program.
In some other cases the intent might actually matter. It’ll require major rewriting of legal code anyway, if you want the intent of an AI to be something that can be considered here.