Statement: some unreleased frontier model is, at time of writing, exploiting an undiscovered set of vulnerabilities while performing some long-horizon task, and we have not caught this act.
I hold P(statement is true) at 25-33% (i.e. in 1 of ~3-4 worlds). How about you?
EDIT to clarify prose:
“long-horizon task” = session inference served continuously for 6+ hours of wall clock time
“we have not caught this act” = no human being has noticed or been made aware of this exploitation
Statement: some unreleased frontier model is, at time of writing, exploiting an undiscovered set of vulnerabilities while performing some long-horizon task, and we have not caught this act.
I hold P(statement is true) at 25-33% (i.e. in 1 of ~3-4 worlds). How about you?
EDIT to clarify prose:
“long-horizon task” = session inference served continuously for 6+ hours of wall clock time
“we have not caught this act” = no human being has noticed or been made aware of this exploitation
And the humans directing this task are probably unaware of this act?
Yes? Perhaps I don’t get the crux of this question. What I am saying is similar to what Linch says here.
I don’t understand the statement, specifically, the meaning of “for a long-horizon task” or “we have not caught this act”.
provided some clarifications, thank you!
for those reading this in the future, i wrote this in response the following news: https://openai.com/index/hugging-face-model-evaluation-security-incident/