PhD student at MIT Sloan trying to pivot to AI safety research. Open for collaboration and discussion. You can reach me via maoc@mit.edu or connect https://www.linkedin.com/in/cmao/
Chengfeng Mao
It is trustworthy because it is, by definition, allied with its principal and shares its values and goals.
I’m unsure how this works, even if we assume the LLM is benevolent and understands the user unusually well. The problem is that human values or desires are not coherent or consistent. We have subconcsious desires, mimetic desires, revealed vs stated preferences, obsessions, aspirations, and etc. They cannot be easily distentangled or ranked and they often conflict. Which should the GA align to? For someone addicted to online sports betting, does the GA help her find better odds or does it treat the betting preference as a local failure mode and help her stop? How about someone obsessed with climbing the corporate ladder?
The GA could act like a guru or life coach, and help you overcome compulsive, parochial, or status driven desires. That is not necessarily bad, and human preferences are often constructed anyway, but this could get murky quickly. For someone growing doubt in a faith, or ending a marriage, at time A they’d want the GA to help them commit; at time B, to help them leave. Which way should GA choose? This also creates a new autonomy problem. If the GA proposes an alternative life direction that seems more meaningful than the user’s current self-understanding, how much authority should that proposal have? Is this empowerment or disempowerment? This is especially important because the interaction would likely be much more influential and omnipresent than ordinary friends or coaches. Self-determination theory seems relevant here because autonomy is not merely getting the option one currently asks for; it is experiencing oneself as the author of one’s action.
All being said, I’m quite sympathetic to this idea and I genuinely hope it would work. I’m a person of low agency and frequently suffer from the conflict between instant gratification vs my higher purpose. I’ve also been trying to build an exobrain like system, which is kinda similar to GA, but much less autonomous. I think for people who have already put in significant amount work in self discovery, this might be helpful.
Thank you for getting back! It’s surprising that there is not much difference in 32B between placebo vs real, given that 32B is better at identifying the drug.
My idea was inspired by Doc-to-LoRA, but I didn’t know HyperSteer, thanks for sharing!
Super interesting work! Sharing a few thoughts:
Given the models are able to introspect the steering vectors at some nontrivial degree, in the free-play setting, is it possible to also show the model all the steering vectors with anonymous names and placeholder as in the prefill experiment, then let the model choose? I think this could more cleanly see if the model has a preference for some “raw feeling” without the confounding from the vector names.
In the redosing experiment, I’m also curious how do the distributions differ in real vs placebo. RIght now it’s just showing the real case.
Another related idea I’ve been thinking is whether it’s possible to train an LLM to learn to output a steering vector given a description about a direction, with some adaptation to the un-embed layer, similar to a hypernetwork. Perhaps then test whether an LLM can design a steering vector which itself will be obsessed with!
Remove anything I feel an unsatiated desire for and become agitated when I don’t have access to. No video games, TV, Netflix, social media. I turned my phone display to greyscale and turned off all those UI animation effects that make interactions feel sleek. I installed https://screenzen.co/ Screen time also works. I asked my partner to set passwords for them. On my laptop, I installed https://selfcontrolapp.com/ and set max block time to months https://gist.github.com/gschema/16bc1e77833dfe06e63b81256473fe72 Whenever I found myself developing any new addiction to a website, like novels on wikisource, I added them to the block list.
Finding a replacement could ease the withdrawal. Lesswrong is a perfect replacement for social media for me. The updates and feedback are much less frequent, and there is no endless scroll.
Perhaps finding a new superstimulus also helped. I got addicted to a mobile game a few weeks ago, and totally shifted my attention from social media to it. When I went cold turkey on all distracting apps, it felt easier because I had less attention on social media. Also, I haven’t developed much attachment to the game given the short time.
I recently realized that self-improvement wouldn’t work without dopamine detoxing (at least for me). I’ve tried all sorts of ADHD medication and self-improvement techniques like journaling, GTD, 5 secs rule, mindfulness, counseling, etc., none of them stuck. I aim to be more self-disciplined by instilling them into my life, but they all require self-discipline to enforce. Especially if I’m surrounded by superstimuli like social media and meme videos, anything that requires attention away from them becomes extra annoying and painful.
When my brain is adapted to superstimuli, it also takes more willpower to shift my attention to something less stimulating or addictive. So I procrastinate more. I wait until the last moment when intense stress can override my craving for dopamine spikes. This also means I gradually lost my agency. I had to wait for external crises to push myself out of indulging in the superstimuli. But self-improvement and contributing to meaningful causes require serious thinking and work, beyond the bare minimum effort to avoid failing at whatever I’m doing. Only recently, after I went cold turkey on all the superstimuli, did I finally feel free again. Doing serious work is no longer that excruciating. I reclaimed mental space for more reflection. The flywheel of self-improvement finally starts spinning.
I think a combination of stated preference parsed by LLM + recommender system trained on revealed preference will likely be a more balanced approach.
Telling the system my preferences probably works well in cases where I’m quite certain, like something I absolutely love or hate. But in a lot of cases, I’m uncertain or don’t even know what I want until I see something and show that through my revealed preference. That was the problem motivating research like this https://www.lesswrong.com/posts/k8SbrC8EMq2RpCmNg/post-mortem-ing-my-earliest-ml-research-paper-7-years-later The bias in revealed preference and stated preference is also a classical direction in the behavioral econ and psychology research: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=437620 https://papers.ssrn.com/sol3/papers.cfm?abstract_id=992869
ML based recommender system seems to be capable of magically learn my preferences that I can’t describe, that’s why it is so addictive. But this could also be beneficial in cases where it can help us discover interests that even ourselves are unaware of.
I’m not that familiar with the latest research in recommender systems, but I’ve noticed a few work that have also been trying to mitigate the tension between addictive instant gratification vs long term well-being: https://arxiv.org/abs/2207.10192 https://arxiv.org/abs/2406.01611 combining LLM into these systems is also very trending.
I’m very susceptible to recommendation algorithms and have been fighting with binge watching videos or social media posts. I recently found that going cold-turkey is much easier than limited usage per day. When I have a quota to check these instant-gratification inducing apps, the temptation is always in the backdrop. Whatever I’m working on feels much more boring compared to checking those apps. Instead, strictly blocking these websites frees me from deciding when to use my quota, and eliminates the comparison. The work I need to do feels less boring. I went cold turkey by asking my friend to set the password of screen time on my phone, and installed self control on my laptop with prolonged intervals of 3 months.
that’s true. standing desk with a mini stepper may also help
I had a similar experience using Claude with paper writing. You summarize my sporadic annoyance really well. A few more minor observations:
It loves making up noun phrases like “prospective rules”, “compilation gaps”, etc. They read legit for someone new to a field, but look awkward and nonsensical to a domain expert.
Such deep-rooted tendency or other preferences like the usage of em dashes are difficult to override with rules or memory files, especially when Claude focuses on other cognitive demanding tasks, or when it hasn’t recalled or applied the relevant rules within the past ~100K tokens.
I mainly use Codex and Gemini to review my paper, as I find them more critical than Claude, especially on Claude’s own writing. Then I ask Claude to address their critique and iterate with this loop. However, this loop seems endless. The reviewers always find new issues in each round, but they never report everything they find. Also, as @lc mentioned, reporting non-issues is common, like over-generalizing beyond the scope I already defined and suggesting caveats at a paranoid level.
Chengfeng Mao’s Shortform
My favorite thing to do when waiting for my AI coding assistant is to do burpee with push-up. Sitting for long hours is bad for health, and taking even a 5-minute break in between could yield some health benefits. Burpee is a great compound exercise to get the heart rate pumped in a few minutes. Squat is also nice. I also noticed more energy and mental clarity since I developed this habit.
Thank you for sharing your experience! This also resonates really well with me. Executive function deficit is also a common symptom of ADHD, and I’ve gone through dozens of self-help books or techniques, like atomic habits, journaling, time blocks, pomodoro, hourly logging etc. None of them stick for more than 3 months, and none have substantially changed how I behave. I recently talked with Claude with this issue, and it points out that the habits I’ve been trying to build up require the exact executive function that I’m aiming to enhance, and this easily breaks when the excitement from novelty fades away. Then I started building a journaling app quite similar to what you describe, with some additional features like syncing with my Outlook calendar and helping me automatically block time and reschedule.
I haven’t started implementing the exo-brain part as you described. I’m quite curious how did you organize your daily logs into long term memory for the LLM so that they fit into the LLM’s context window? Do you follow any architecture from relevant research?
Thank you for the response!
I like this idea! Perhaps the GA can also help people who struggle on the same things to connect and help best practices to diffuse faster. there could be a chance of doing life logging and self improvement at scale, and more reliably identifying effective interventions faster. One tricky thing is the trade off between privacy vs the how well the system can learn.
If you are building this, I would be happy to learn more or help. I’m also building a system to monitor my computer us e activities, automatically scan my calendar, and help me decompose and prioritize tasks. I’m also trying to make it more proactive to help me align my sit short term actions better with my long term goals.