In my world the typical time is more like 5 weeks or months :)
Tapatakt
I currently think this is too de-anonymizing
Ok. But no need to show exact information if the point is to distinguish karma +950 −1000 from karma +0 −50.
Maybe display approximate ranges of upvotes/downvotes? Kinda like creature quantities in HoMM. “Swarm (250-499) liked, horde (50-99) disliked”.
Or even just show what fraction of overall karma was in other direction rounded to 10% (or 20%, or even 25%).
introduces some kind of weird social dynamics that I think would be on the margin worse
Would be moderately interested in elaboration.
But then someone can screenshot the call to violence with the karmameter, which is good PR.
I think “there are bad people in our community, but the overwhelming majority of us don’t like what they say” isn’t usually used as PR. I suspect it’s because it wouldn’t actually work.
Maybe I am failing to parse this, but I much prefer for calls to violence in other communities towards me to also face clear pushback, and to be public (since then I can orient to them and take precautions), so I don’t really see the prisoner’s dilemma here. Indeed people threatened by the potential violence are the people who most would like to know about the threat!
I’m not sure about this in any direction. But it seemed missing from Samin’s point of view.
Thanks for starting the discussion!
My thoughts:
“Displaying the number of or karma from upvotes or downvotes on hover instead of just the absolute number of votes” is good idea anyway
Additional for allowing:
Slippery slope into censorship.
Litany of Tarski. If violence works, etc.
Additional against allowing:
Someone can screenshot the call to violence without karmameter, bad PR.
Epistemic Prisoner’s Dilemma. We would want people who think something like “If no one builds it, everyone dies” to forbid calls to violence against pro-pause side in their communities.
Continuation of this trend already requires some form of TAI. The method of how AI systems generate value has to radically change. Otherwise who would pay so much money for them?
It’s kinda like making similar argument about parameter numbers and saying “and if it’s more than … parameters, it means Earth surface is all computronium, so obviously AGI was achieved”.
I already mostly believe in the logical implication “no AGI → break of trends”, so “no break of trends → AGI” is not an additional argument.
I vibecoded a plugin, tested with Brave.
Though of course this particular example does not prove that the game of Catan, in particular, has situations like this.
A has 7 points, “Year of Plenty” card, 1 brick and 3 wood. A can get Longest Road either by building 4 roads or by breaking B’s road with a settlement, but to build this settlement A has to first build one road.
B has 9 points including 2 points from Longest Road and enough resources so they can build a settlement in one turn unless 7 is rolled.
C has 9 points, 1 brick and can maybe win in one turn depending on dice rolls.
A’s turn.
A: “I have Road Builder card, 3 wood, but only 1 brick. C, can you sell me brick for wood? I will build 4 roads, get Longest Road, B will not win in their turn and then we both have a chance.”
C: “I don’t really need wood, but I see that B probably wins if we don’t do it, so OK.”
A plays “Year of Plenty”, takes grain and wool, builds one road and a settlement, wins the game.
I would add 6, but not 9.
The list of frontpage posts and all opened posts appear in fake windows that can be drag-dropped and can obscure each other.
My thoughts about this:
1. Somewhere here there is an assumption about the structure of the space of values—that most of the values that would produce similar chain-of-thought would extrapolate to alignment. Maybe it is like this, and if it’s not that would mean that probably if we increase the intelligence of some really very good person to superintelligent level it would still have catastrophic consequences to everyone else. And in general without it alignment is probably doomed anyway. I think if we have to make one assumption on basis “if it’s false, we are doomed anyway”, this one is not the worst, but it should be explicitly labeled like this to avoid part ways with reality completely by making a lot of such assumptions and not only one. …actually even if this assumption is true for humans it doesn’t mean it is definitely true for LLMs, because they are not humans.
2. Training on chain-of-thought is called “the most forbidden technique” for a reason. Using chain-of-thought to select what model to expand/upgrade/use its outputs to train other models is not exactly “training on chain-of-thought”, but it’s close. How many bits of selection pressure would it applied? How many bits of selection pressure probably can be applied without making chain-of-thought untrustworthy? Which of these two numbers is greater? How sure are we about it?
>the most urgent film of our time
>look inside
>only in theaters March 27
Ok, guys, really, does anyone (Claude says probably not) track if there are negative utilitarians in leadership of top AI companies?
That’s kinda important, don’t you think?
Happy New Year, btw.
UPD: Obviously people think it’s not a good point. Why? Do you think it’s not important, not neglected, or that answer is obviously “no”?
If a continuous function goes from value A to value B, it must pass through every value in between. In other words, tipping points must necessarily exist.
I propose more specific idea: if you are uniformly uncertain about fractional part of , then .
E.g., if you hurry on the way to the subway station without knowing when the next train arrives and got there 10 seconds earlier than if you didn’t hurry, you win exactly the same 10 seconds in expectation.
Untestability: you cannot safely experiment on near-ASI (I mean, you can, but you’re not guaranteed not to cross the threshold into the danger zone, and the authors believe that anything you can learn from before won’t be too useful).
I think “won’t be too useful” is kinda misleading. Point is more like “it’s at least as difficult as launching a rocket into space without good theory about how gravity works and what the space is”. Early tests and experiments are useful! They can help you with the theory! You just want to be completely sure that you are not in your test rocket yourself.
At times the authors appeal to prominent figures as evidence that the danger is widely acknowledged. At other times, the book paints the entire ML and AI safety ecosystem as naive, reckless, or intellectually unserious.
I see no contradiction between these two statements:
Prominent figures and also median experts believe that the risks are at the level we can surely call totally unacceptable (even if some experts themselves consider it acceptable)
Current field of AI research can’t make much progress on AI alignment problem.
People totally can know about the risk without also knowing what to do about it.
Thanks for your concern!
I think I worded it poorly. I think it is an “internally visible mental phenomena” for me. I do know how it feels and have some access to this thing. It’s different from hyperstition and different from “white doublethink”/”gamification of hyperstition”. It’s easy enough to summon it on command and check, yeah, it’s that thing. It’s the thing that helps to jump in a lake from a 7-meters cliff, that helps to get up from a very comfy bed, that sometimes helps to overcome social anxiety. But I didn’t generalise from these examples to one unified concept before.
And in the cases where I sometimes do it, my skill issues are due to the fact that the access is not easy enough:
I can’t do it constantly, it takes several seconds and eats attention.
I can’t reliably remember to do when it’s most important—in highly stressful situations or when my attention is too occupied with other stuff.
Some internal processes (usually—strong negative emotions) can override it by uploading more powerful image into the script, so I follow it instead, even while understanding that it’s worse.
Also it doesn’t really work for long period of time from one uploading. (So it works best when returning to default course of action after initial decision would be hard/impossible/obviously silly/embarassing/weird.)
Do you think I’m wrong and this is a different thing?
Thank you! Datapoint: I think at least some parts of this can be useful for me personally.
Somehat connected to the first part, one of the most “internal-memetic” moments from “Project: Lawful” for me is this short exchange between Keltham and Maillol:
“For that Matter, what is the Governance budget?”
“Don’t panic. Nobody knows.”
“Why exactly should I not panic?”
“Because it won’t actually help.”
“Very sensible.”
If evil and not very smart bureaucrat understands it, I can too :)
Third part is the most interesting. It makes perfect sense, but I have no easy-to-access perception of this thing. Will try to do something with this skill issue. Also, “internal script / pseudo-predictive sort-of-world-model that instead connects to motor output” looks like the thing that has a 3-syllable max word about it in Baseline. Do you know a good term for it?
However, I feel that all this is much more applicable to the kinds of “going insane” which look like “person does stupid and dramatic things” and less (but nonzero) applicable to other kinds, e.g., anxiety, depression or passive despair at the background (like nonverbalized “meh, it doesn’t really matter what I do, so I can work a little less today”).
list of fiction genres encompassed by almost any randomly selected… say, twenty… non-“traditional roleplaying game” “TTRPGs”.
Hmmm… “Almost any genre ever” for Fate? (Ok, not the genres where main characters must be very incompetent.) I personally prefer systems with more narrow focus which support the tropes of the specific genre, but your statement is just false.
D&D is good for heroic fantasy and mixes of heroic fantasy with some other staff. D&D is bad for almost everything else. Of course, some modules try to do something else with D&D, but they usually would be better with some other system.
Random thought: maybe it makes sense to allow mostly-LLM-generated posts if the full prompt is provided (maybe itself in collapsible section). Not sure.
Obviously, there are situations when Alice couldn’t just buy the same thing on her own. But besides that, plausible deniability:
No one except Bob knows the exact money and attention costs of a gift and how exactly they compare with his gifts to other people.
No one except Alice knows exactly how much she likes the gift, incuding when comparing with gifts from other people.
Absolutely no one knows how both previous points compare between Bob’s gift to Alice and Alice’s gift to Bob.
No one knows if Alice would buy the gift on her own of she had this idea, so no one can critique her for wasting money. Bob has a free pass, because it was a gift, he was altruistic.
Would you also approve other costly signals? Like, I dunno, cutting off a phalanx from a pinky when entering a relationship.
I think this post is very underrated. I often think about connections and differences between “CDT with anthropic uncertainty about ghosts” and FDT/UDT. They don’t seem to actually be equal, but I think there is something interesting there.