That’s the hope!
Lukas Petersson
Also, the first step to do anything defensive about it. Evals are for sure dual use, but I’m leaning towards better to know than to stick your head in the sand.
Personally, I’m more worried about loss of control than misuse, which in my mind makes Chinese vs American capabilities less relevant (please enlighten me if I’m missing something here tho, maybe I misunderstood your question).
The alternative is to let the AIs take actions like “go forward 3m”, “rotate 25 degrees”, etc at every iteration. But we found that 1) this makes the movement way too slow to follow anyone and 2) models are really bad at this.
Should we be worried about how good AI is getting at coding autonomous drones?
Towards A Happy Future With AI Employers
Should LLMs accept invites to Epstein’s island?
I see. Thank you!
Hi again, should I assume it’s not happening?
RT-2 (the paper you cited) is a VLA, not LLM. VLAs are what the “executor” in our diagram uses.
LLM robots can’t pass butter (and they are having an existential crisis about it)
Hey Ted! Any updates? :)
We set it to some date in the future
Thanks! Vending-Bench v2 is going to be fire. Would love to include gpt5 <3
This is a great point. I admit I have to better understand what each model provider does behind the scenes in the API. Sad if the days of access to the model is gone.
We thought about that, but then it’s not reproducible if we want to run it for new models later
Thanks, that would be great!
You can’t eval GPT5 anymore
AI misbehaviour in the wild from Andon Labs’ Safety Report
Thanks for highlighting our work!
You wouldn’t believe how bad they are at using controlers of that format.