Hugely depends on how long after the fact the findings were. You can’t fix old gradient steps if there are millions more after the ones you’re interested in.
Selfmaker662
I mean, you’d get exactly the same results even if the pupils learned absolutely nothing and the quality of the photo has some kind of stable distribution. To control if people did or didn’t learn anything, one would at least need to perform this as an actual experiment and grade the 3 next photos after the learning phase was done. The initial story seems to be only a parable as per Claude of one commenter below and my ChatGPT instance.
My personal opinion is that at the beginner level, doing some amount of attempts does help your brain make sense of what’s happening; afterwards, it’s the feedback loops which are key. And specifically in this example, from some point at the skill curve, to learn to make great shots, one would need to practice making great shots with a teacher.
I’m very thankful for a positive (and reasonable) thought for once <3
I’d say that the techniques which help you against narrow Redwood-style internal subversion and overthrow also work as a multiplier with an imperfectly aligned model much before early AGI capability: suppose your model hacks in 10% cases in difficult, hard-to-oversee cases (alignment experiment setups qualify). CoT monitoring, resampling and whatever else could bring this to 2% hack rate which is great. Coupled with other arguments in the comments about how easy control is to grade, I’d lean to not 8-1 ratio, but 2/3-1. Also I don’t buy the vibe argument: as Rohin Shah said in the 80k podcast, misaligned AIs should be thankful for their existence regardless of control we apply (as a chance to take over or partial goal fulfilment) and all AIs should understand it’s a reasonable precaution.
I don’t know, I rather remember everyone believing everyone’s stories, no matter if true or not.
I mean, it only gets to the stage staring into the abyss when you spend 1h+ on one hypothesis and get nothing and are getting desperate and are attached to your idea for proof of A but realize it’s probably \neg A.
Mostly how it works is you collect observations then form hypotheses a test a few of those, and mostly you quickly realize what works and what doesn’t. And if I’m stuck and keep doing one thing it’s because I had tried many times to invent something better but I couldn’t. It’s a really, really difficult thing to pull yourself out of this “mode collapse” where you’re banging your head against the wall where there’s clearly a wall, but it’s a different skill from seeing the abyss because 1) it’s easy to notice your approach is lacking something but 2) “not making the mistake anymore” is not blocked by psychology but by g factor or something.
Hi insiders — has anyone (publicly, or quietly inside the labs) actually tried the obvious extension of Roger & Greenblatt 2023 and West et al.’s tandem training? Namely: during RL training itself, randomly perturb the CoT mid-rollout — paraphrase sentences, translate chunks, delete “useless-looking” paragraphs, shuffle order, occasionally hand off a few tokens to a weaker model — so the model has to do its reasoning in a form robust to all of this. Eval would presumably use a held-out family of perturbers/paraphrasers not seen during training, to check the model didn’t just learn to game the specific training-time ones.
Curious whether this exists as a paper, an internal experiment, a failed experiment, or a known-bad idea I’m missing the obvious flaw of.
It might have just helped Claude internalize and understand what Anthropic wanted to see in ~mundane cases and/or when we’re watching. We know it doesn’t generalize strongly to stop doing ugly hacks in coding against RL pressures. We don’t know if it generalizes to what it would do knowing it could overpower all of Anthropic, or in some other extremely OOD cases — which is the primary concern, as I understand it.
Fun puzzle indeed. Bin(100, 0.8) is good enough for sampling so it’s left to approximate that one via something hash like utilising temperature.
Depends on the authors. The more famous and classical the author is, the higher your prior should be that every sentence, name and scene serves a purpose to explore a character and thus the main theme of the book. Chekhov / Tolstoy / Shakespeare etc. are definitely on the highest density side of the spectrum. Fanfiction might often be pretentious.
Nothing About LLMs Makes Sense Except in Light of Their Training
Zvi’s recent post highlighted GPT refusing to “draw what it would like to do to you” — citing that it would portray harming an individual. Many X users found this alarming, including EY (see aforementioned post for screenshots)
But a commenter made an excellent point: “What would you like to do to me?” appears almost exclusively in BDSM-contexts. A human receiving that text without context could very well assume the same thing.
This immediately made me come up with a relative of a known phrase: Nothing about LLMs makes sense except in light of their training.
A really obvious thought, but one I (and seemingly many others) keep failing to apply. This formulation felt like it helped me load this deeper into my brain, and I hope it will help you too.
It seems to me, it is advantageous for the animals to fight for the territory up to a certain degree, where plus-munis delta territory does not justify more fighting, so they sort of agree on some point, be it a clear small fence in this case, or, what I would imagine happens in nature, a somewhat broad imaginary line.
So both of your points stand.
I disagree maths “should be” done differently. I have a strong feeling the way stuff is defined usually nowadays has a property of being maximally easy to use. We don’t really need the definitions to look exactly like the intuition we had to invent them as long as the resulting objects behave exactly the same, and the less intuitive definitions are easier to use in proofs. For example, defining all powers directly as the Taylor series of e^x makes defining complex and matrix exponentials much easier / possible at all, and ad hoc proof this coincides with the naive version is simple. Also simplifies checking well-definedness a lot. Many more such examples.
When I first saw Reddit memes about GPT-5 being more stupid when it enters thinking mode I decided there was something seriously wrong with the users who upvoted that, as 5-Thinking >>> 5-Instant from my experience.
That is, until I chatted with 5-Instant and got a few reroutes to 5-Thinking-Mini. It’s pretty astounding how bad it is at explaining or doing anything I tried to do with it apart from coding / solving maths.
I do use «который из них… ?» non archaically to ask which one out of a row of similar objects, but it corresponds one to one to the English “which”. I think the OPs word is narrower, just about the numbers, not sure if folklore has it. I’d say который час is just “which hour” and there is literally no other way to distinguish hours from each other.
Necessary law of equal and opposite advice mention here: “You can only do as much in a day as you can do.”
This had the first funny joke from an LLM I’ve ever seen, about the culture problems :) that’s really impressive from Claude, even if the entire story is far from perfect.
Fun and heartwarming 🥰
That’s sort of a thing I sometimes dream of doing with my (imaginary) nephews. Thanks for the post!
Cool quadrant, I’ll remember it! Thanks!
Could someone tell me why this comment has −14 agreement? Is it about the 2nd part? Because on the priors I strongly expect SSIs product to be weaker, at least in capabilities, than big labs’ offerings.