Isn’t that just the trolley problem with chickens? Absent context I value five humans more than one human, that doesn’t necessarily mean I would murder one human to save five. I don’t see why saying that chicken suffering is bad necessarily traps you or obligates you to kill humans to protect even a very large number of chickens.
Altaer
Out of curiosity, how relatively bad do you think human and chicken suffering is?
I also think a chicken being tortured is much less bad than a human being tortured, but I wouldn’t rush to weighting a chicken’s torment at, say, less than 1/1000th of a human’s. That lower bound gets me to the naive conclusion that 4.5 billion chickens experiencing some form of suffering is likely at least as bad as 4.5 million humans experiencing something analogous. From that perspective the comparisons in the post seem pretty reasonable to me.
Edit: Follow up question—do you actually think the human/chicken suffering comparison is irrational or confused, in addition to being symptomatic of the author’s bubble?
Yes, sorry, I understood the point of the post. My confusion was more about why anyone would find the intuition credible to begin with, even prior to this hack occurring.
It seems like the initial intuition can work only if the model is able to convince me that it can and will cheat on its task, but is choosing to negotiate with me instead. But if this happens why wouldn’t I just shut down/stop using the model?
Is the idea to train models that essentially choose to say “hey I would successfully reward hack the task you’ve given me, therefore I won’t do it”? It seems very difficult to train a model that would do this, since if gets rewarded for admitting that it would cheat it would have no incentive not to always do so, whether or not it’s actually true.
It would help me if you gave an example of when a model would benefit by choosing to ‘negotiate for its reward’.
I see—I wonder how sensitive this measurement is to labs doing post-training steps which include lots of low/no reasoning token math problems. With K3 currently supporting only ‘max’ reasoning I’m also led to wonder if if Kimi deliberately put lower focus on low/no-reasoning tasks in their post-training and if you could in fact be seeing the results of that rather than of a substantially weak pre-training phase.
I’m curious, what tests are you running to isolate the quality of the K3 pretrain? If I was trying to do this I’d have pretty low confidence in my ability to correctly credit the model’s pre-training vs post-training for its performance on any given test or suite of tests.
I guess I can take a stab: the artist’s use of the “marriage” and “spouse” verbiage is probably the weakest evidence in the post for the conclusion that the relationship with the dog is sexual, with the more compelling evidence including Gossiaux talking about french kissing the dog, describing an intention to portray the dog in a way that’s open to the interpretation of sexuality, describing trying to depict the dog’s “wet dream” etc. So picking on the points you did and calling the post’s conclusion a stretch might read like attacking a strawman—the conclusion isn’t actually based on those (weak) points.
Altaer’s Shortform
For anyone doing empirical alignment research, could you share some key tooling/systems-level challenges you’ve run into? I’m considering ways I could contribute to this part of the stack and I think any real world experiences people could share would help me come up with a more grounded approach. An example of the kind of response I’m looking for: “Pytorch doesn’t support the type of weight-inspection I need to do in a performant enough way/with my model-parallelism setup.”
By extension would you also have no preference between killing a chicken and killing a potato plant?