LessWrong dev & admin as of July 5th, 2022.
RobertM
@beyarkay (Boyd Kane) there was a broken image pointing at (I believe) a localhost url underneath the “processed certain configuration files for uploaded datasets:” line; I’ve deleted it since it was causing Chrome (and maybe other browsers) to ask users for permission to “Access other apps and services on this device”. (And I’ve filed that in our bugs channel.)
You may want to re-upload the image and re-publish the post.
For future reference, I normally wouldn’t approve this kind of comment from a new user (an outbound link to substantially AI-written content), but permitted it for the sake of the essay contest, since it was disclosed. I may change my mind about this kind of thing in the future if it becomes a problem.
He could have “meant” that they’d paused training of that specific model. I would be astonished if they’d at any point stopped running all code that would result in updates to any model’s weights.
Error: NotFoundError
The error in the screenshot above looks like a buggy interaction between Chrome’s auto-translate feature and React.
app.operation_not_allowed
See my comment here.
I think they are slightly ashamed of censoring things, and it doesn’t pull their enthusiastic attention and desire to make it work well, and so it doesn’t get the dev and PM efforts that would go into, for example, an April Fools event?
No, not particularly. It is true that working on improving the user experience here is not especially motivating, but this cluster of bugs + bad UX is near the top of my internal to-do list of relatively important things to work on; there are just a lot of things to do (and that list seems to be growing rather than shrinking).
(Like the font isn’t even the right size! It is the only thing that matters on that page, and it is so tiny as to seem an afterthought… except that it is red so they clearly don’t intend it to be an afterthought. The programmer who set the color was thinking it should be big and obvious, and the programmer who set the font size had a different intent.)
The error message color and font size is one of the few remaining artifacts of the original framework that LessWrong 2.0 was built on (VulcanJS). Our story for correctly surfacing legible errors to users is quite bad, across the codebase.
And I think it is probably a bug to not let changes be published? (Though maybe they are afraid of some kind of adversarial dynamic and this is correct in order to make things hardened?
I made this change because I wanted to prevent people from making the contents of the rejected posts displayed on lesswrong.com/moderation misleading (with respect to the actual content that caused them to be rejected). The rejection feature was not designed with “post is modified to be un-rejected” in mind; this is not well-communicated in the UI. You should simply make a new post.
But it is at least it is a bug to accept edits (which consume time) and then refuse to let them be published (with no warning or explanation of this)?
Yes, this is basically an oversight.
I guess in a deeper sense, maybe their own policy is not actually be well articulated, plausibly because it was designed by a committee to satisfice the not-perfectly-compatible desires of various stakeholders, some of whom likely had incompatible mental models?
The policy about what users should (and should not) do is clearly described in the post that Seth linked above. There is no English-language description of all of its downstream technical consequences because we don’t have the (truly absurd) bandwidth that would be needed to satisfy that requirement in full generality, across all of the site rules (and other things that might motivate moderator action).
The policy was “designed” by me marinating in the ways in which the previous policy was inadequate over the course of many months of moderation work, writing up a new policy, running it by Habryka, adjusting the wording slightly, and then publishing that post. LessWrong sees regular engineering and moderation contributions from 6-7 people, and usually only 2-3 people on any given week. We do not have the people to form a committee.
I do not see any emails from you to team@lesswrong.com, which should be automatically forwarded to our Intercom inbox. What email address did you send the bug reports to?
So https://openai.com/index/safety-alignment-long-horizon-models/ was published yesterday. The NanoGPT speedrun PR was put up before May 14th (because PR 287 was referenced by PR 300, which was put on May 14th). OpenAI presumably did all of the things they said in that blog post (paused deployment, improved alignment, additional guardrails, etc) in the meantime.
Then, some time in the last 1-2 weeks, the same(?)[1] model breaches Hugging Face. Hugging Face publishes their blog post on July 16th. https://openai.com/index/hugging-face-model-evaluation-security-incident/ is published today.
This timing is kind of surprising! Maybe suggests that the arm of OpenAI responsible for publishing the first blog post did not know about OpenAI’s responsibility w.r.t. the HF incident yesterday. As it is, all of the claims about the additional work done seem a little… something.- ^
Not actually specified by the blog post; I suppose it could be a different model entirely? But it does seem like a simpler explanation that it’s the ~same model, or at least within the same model lineage, than that OpenAI has 2 different unreleased highly cyber-capable models.
- ^
Yes, I agree the government will care a lot about this. The entire point my post is trying to make is that this will be a hard problem for the government to succeed at, and then have any justifiable confidence in their success. Do you have any sense of what I could have written that would have made that clearer to you?
The government could check if models are useful to them by directing automated researchers to create a new type of AI with values based on defined specifications for military purposes, subsequently running it in isolated, high-resolution military simulations to test hundreds of thousands of scenarios.
I think you are imagining a world very different from the one I’m imagining here, but not totally sure.
I think this was explicitly covered in the post:
It’d be really dumb to kick off RSI and then :shocked_pikachu: if the USG put a gun to your head and told you “Hey, buddy, that goal slot you got there? We’re putting our goals in there.”
(With that said, I’m not totally sure I understand whether you’re disagreeing with something in the post, or making a different point.)
The US Government may find it difficult to seize control during takeoff
If you establish the integration, and then click the actual “Claude” button in the publishing menu, it prefills the claude.ai prompt with this message:
I’m writing a post on LessWrong and would appreciate your inline feedback on it. The post is at https://www.lesswrong.com/editPost?postId=[redacted]&key=[redacted] and documentation for interacting with the site’s API is at https://www.lesswrong.com/api/SKILL.md.
I thought that’s what you meant by this:
I also tried the approach described here of setting the post to allow comments to anyone with the URL, and Sonnet retrieved the content but considered the associated API notes to be highly suspicious and again, likely a prompt injection attack on me.
But now I realize you probably meant the metadata we return as part of the markdown response when agents fetch posts via the common mechanisms that we can detect.
They keep getting more and more suspicious with each model generation. (Though I haven’t heard of Fable complaining yet...)
Re: the SKILL.md file—you can just go to the page yourself and copy its contents into the context window, if they refuse to fetch them.
Why not an official “Connector” using an MCP server? A few reasons:
MCP-based tools are (currently) non-composable, so in many ways are less ergonomic for models, i.e. they can’t pipe their outputs to disk, into other commands, etc.
Anthropic’s approval process is annoying.
For use in contexts like Claude Code (or, really, anything other than the claude.ai web interface), it’s almost strictly more overhead to set up than to create a new skill.
We’ve iterated on the wording and will probably continue to do so, but it’s a moving target.
This is somewhat more LLM-edited than ideal. Please see our current policy on LLM use for future reference. (I currently have no idea whether this is a meaningful experiment to have run. I haven’t read the original paper and maybe the entire thing is vacuous? Who knows.)
This seems like a plausibly interesting experiment to have run (though the results do seem pretty ambiguous), but the text is somewhat more LLM-edited than ideal. Please see our current policy on LLM use for future reference.
I assume you are having a bad time but LessWrong is not the place for doing it in quite this way. I’ve removed your commenting permissions. Feel free to DM me if you would like to explain what happened here.
Oh, I have a guess about where the miscommunication is. I took Linch’s “I do not have a clear understanding of why this occurs.” sentence to refer to the phenomenon in general, not to the headline claim of unreleased models being substantially more reward-hacky/locally deceptive/etc than released models. And I don’t need hard data to personally observe that the released models are noticeably reward-hacky, because I have personally noticed.
Oh, I see. I guess that’s reasonable, but, eh, I feel like it can mostly be screened off by the personal experience?
Mod note: this reads pretty LLM-y to me. Please avoid using LLMs for stylistic edits in the future.
I admit that I don’t quite understand how this comment is responsive to that particular line. Are you trying to dispute the underlying premise that both Linch and I accept as true: that the models do in fact frequently engage in reward hacking? Because I have plenty of personal experience with recent generations of frontier models engaging in behaviors that are reasonably described that way in a deployment setting—I’m not running artifically-constrained evals, I’m actually using them to write code.
Curated. The argument here is straightforward and seems basically correct to me. I think I have shorter timelines than Tsvi does, at least if not conditioning on a successful pause/stop effort, but as Tsvi says:
Really, most of the HIA has substantial impact even with short timelines section is important to understand, and should carry the argument even for people with pretty confidently short timelines. (Unless they disagree with “Very short timelines are pretty intractable.” by way of things that can meaningfully be influenced by humans today.) The points at which people think marginal effort allocation to HIA stops making sense might differ, but right now, there need to be more people working on this.