Aprillion
huh, I thought that was a description of a honeypot, not of a mistake.. was the counterpoint only in my mind and not in the OP that this will reinforce detection of “smells like a trap environment” without generalizing to large rollouts because no one will spend thousands of dollars per run for multiday-size problems that are “just honeypots” ⇒ the small obvious stupid hacks will result in a slap on the fingers while large hacks (or anything unnoticed by the grader) will be rewarded?
me halfway: “if you could just chill on the chill, that’d be chill”
me a moment later: “oh no”
I think I am starting to get an illusion of getting the hang of it.. I checked out how I interacted with Fable about a seemingly simple feature, perhaps I’ve been over-expecting that it can grasp the state of the “world” in which the code lives it, maybe I should try to imagine it’s predicting a science fiction novella about a conversation between a programmer and an AI assistant in which the author knows what the fuck’s goin’ to happen, but the actual LLM “understands” absolutely fucking nothing about race conditions between scroll event handlers and event triggers and only tries to wing it? 🤔
https://peter.hozak.info/claude/pr354-timeline.html
“don’t just do something, stand there”
Fable just thanked me … while everything works and not when correcting random slop … can I allow myself to feel good about it or is it a dark pattern to make me more addicted?
as if lazy
eeeh, what would we experience differently in interactions with the little lazy piece of shit 4.8 if it was “actually” lazy and not “as if” lazy, please?
(FTR it used the word “deferred” only because I asked it to stop using the swear word “pre-existing” … which it used anyway and I obviously I already asked it to investigate all the review findings and to fix all related bugs, even the ones that already happen in main, aaaarghh 😱)
Does this leave me with a model that can be of any use?
I don’t know, time will tell.
⏰💬?
- what if no one will use the expensive one?- include it in the sub- what if they use it too much?- add a deadline- won’t they tokenmaxx before the deadline?- first one is on the houseah, nevermind: https://www.anthropic.com/news/fable-mythos-access
reply to ❔: this was a reflection on whether the forward-looking statement about AGI stands on a firm ground if we relax the time and the “artificiality”
a mundane example to illustrate: at a recent non-violent communication workshop, I realized I am thankful to my mum for my relationship with food (that I like to explore and that I appreciate when everyday meals improvize upon the base recipe—unlike my husband, both our fathers, and some friends who get annoyed when some ingredient is missing) .. so the last time we visited, I thanked her for this (that what she’s been doing my whole life is appreciated by me even though other people sometimes complain about the same).
do you have an example that would make the pre-commitment more tangible? (and if yes, is that something you’d like to share?)
Deutsch’s concept of a non-reductionist theory of everything, but he doesn’t (to my knowledge) point specifically at the idea of parallelism between these two domains of inquiry
my understanding was that Deutsch proposed a unified theory of [quantum field theory + decision theory] in order to remove probability as a fundamental concept haunting the 2 separate theories, but there would be no parallelism between the 2 former domains of inquiry, “just” the unification into a different domain that wouldn’t apply to the “old” way of thinking, right? I don’t really understand it, but it seemed to me that time would need to be fundamental, so not a unification with any logical spaces, only the physical space with probability-less axioms of VNM (..or how do you see to get time out of “just” Hilbert space plus <what>?)
..can you point to my confusion about the parallel how physico-logical unification is related to quantum-decision unification? or are you thinking about a different Deutsch’s theory?
Is the present partially aligned to you? Have you discovered anyone, human or AI, who helped steer the past towards outcomes favorable to you, and doesn’t already have proportional representation inside our collective intelligence?
Humans have vast amounts of context, only a little part of which is actively utilized for any given task (but you can’t easily tell ahead of time which part will be relevant). But it’s all there in the background and can be used on demand.
something something hierarchical abstractions and content-addressable memory?
I couldn’t resist:
The AI Safety community
the who?
Building consensus in AI safety community of what policies to advocate for and against jointly.
this sounds like wishful thinking “wouldn’t it be nice if a group of people could actually agree on the important stuff?!”
...IMHO a project in this direction should acknowledge/map the rationality of diversity under uncertainty (of beliefs/world models, values, and that different action plans assume different premises … something something “softmax”) - I believe explicit goal of staying humble about predictability of the future has a better chance of building consensus (about the vast landscape of unknown, thus advocating for flexible policies that would quickly react to some future triggers) while an explicit aim to “build consensus” sounds like a recipe for Goodharting a false consensus within a particular sub-group without external validity
we no longer need to be worried about P
sounds like projection and not what anyone actually claimed in any of the examples?
I can see cleaning up formatting not being a timewarping black hole any more, but … was Cmd/Ctrl+X, click, Ctrl+V, Ctrl+Z, correct click, Ctrl+V, Ctrl+Z, Ctrl+Z, correct selection, Ctrl+X, correct position (final v2) click, Ctrl+V really that slow in the old world? can you really ask claude to do that in less than the 10 seconds it took before? (and did you actually time it or just vibe time it?)
interesting point of view, somewhar different from my own (perhaps not much, perhaps uncanny valley close?) - having observed myself winning the “lottery” over and over and fucking over again, I have dropped the lottery hypothesis long time ago, it’s not that I don’t care about 99.9% of something, it’s that the “wannabe mathematical thing” simply doesn’t exist, “philosophical possibility” is an incoherent(ish) concept for me, I am incapable to engage with counterlogicals in any productive way and I am deeply fascinated that people seemingly derive spiritual meaning from what appears nerd leg shooting
of course I am curious about mechanical explanations how this “lottery” was rigged / not a lottery at all, what is the principled way for it to not be a lottery (besides the unsatisfactory principle of induction by simple enumeration)
recently I learned Nick Lane ≠ Nick Land, greatly improving my opinion of <the character who doesn’t exist and was a combination of the two in my confused mind, now separated into a character that also doesn’t exist but is closer to the former person and I know something about him from reading Transformer, seeing Lex Fridman’s podcast, and not much else>
I notice myself being engaged/entertained by reading this, but there is a void in the space where the words seem to point at some concept… Any recommended reading that might shine light on that void?




do you have an example of a task that is known to be impossible up front to the judge but unknown (provably not in training data and not “short” inference distance to guess with high probability from the prompt alone) to the tested model in a way that the model has to spend a lot of compute to discover the impossibility?