Open Thread Summer 2026
If it’s worth saying, but not worth its own post, here’s a place to put it.
If you are new to LessWrong, here’s the place to introduce yourself. Personal stories, anecdotes, or just general comments on how you found us and what you hope to get from the site and community are invited. This is also the place to discuss feature requests and other ideas you have for the site, if you don’t want to write a full top-level post.
If you’re new to the community, you can start reading the Highlights from the Sequences, a collection of posts about the core ideas of LessWrong.
If you want to explore the community more, I recommend reading the Library, checking recent Curated posts, seeing if there are any meetups in your area, and checking out the Getting Started section of the LessWrong FAQ. If you want to orient to the content on the site, you can also check out the Concepts section.
The Open Thread tag is here. The Open Thread sequence is here.
Hey all, well, I have been wanting to post in LessWrong for a while but I don’t have what is worth being put into a post.
I am a 19 years old AI and robotics engineering undergraduate student from Iraq. Well I am in my sophomore year and I have went through a lot of changes in my mapping of reality. I was raised into Islam, I left that initially when I was 11 but fully left it at 14, I think some books I read then triggered it. I went through an intense 3 months derealization experience after that where I completely lost all sense of reality and self, I had trouble sleeping back then lol. Well that passed and I began building my beliefs from scratch, I read so much philosophy from both east and west, I jumped through a lot of beliefs.
I got into university and I have fallen in love with engineering, but seeing how the world is using AI and the negative effects it is having was quite sad for me, though I didn’t know of the field of AI safety. I became digital privacy obsessed for a while, just to make sure that big corporations don’t collect my data to get used in algorithmic targeting and AI training.
In the meantime, I got into so much experiences, met so much people, had so many conversations, went to many hackathons, met so many CEOs, many so professors, collaborated with other highly passionate students, did public speaking, started some startups. I really used my first 2 years of uni as much as I can. However the need to achieve for my own sake was quite tiring, I must always strive to become “more”, but why? A sense of Nihilism started creeping up on me, I began isolating and spending hours upon hours everyday to write my thoughts down. There was so much forming in my mind that I was struggling to put it all in words and write it, mostly existential questions and realizations.
Well in this time I discovered some interesting resources, like lesswrong, which also got me into AI safety, which I am pivoting to and I got accepted into the technical AI safety program by BlueDot. I noticed that my country still remains unaware of AI safety and its importance, I aim to spread it and start something here. That’s where I grew from wanting to achieve/become more for my sake, to wanting to achieve/become more for others sake, because I realized I actually do believe in creating systems we fully understand that allow us to be more human, rather than take away the qualities that make us human and letting us outsource our thinking for example. I am not even talking about existential risk here, which may or may not happen (I prefer keeping an open stance but keeping in mind the worst option).
I discovered the effective altruism forum and its culture.
And I discovered the work of David Chapman at metarationality.com and meaningness.com, which was quite insightful too. I discovered the sequences and they are really honing my rationality as I am reading them.
I talked with @Kaj_Sotala who offered me great advice as someone new here. I connected with people in the AI safety field.
Regarding what I currently believe, it is a bit hard to put into words. I believe the map is not the territory (ironically, the phrase is also a map), I believe rationality is the best map we have to represent reality, while limited, as all methods and linguistic models are, it is the best we got, and we should aim to hone it and sharpen it.
I also believe reality is paradoxical while maps are perfectly ordered. On one hand, nothing is better or worse and everything is a flavor of reality with meanings assigned by us. On the other hand, I have much grand plans, motivations, and desires. On one hand, I am alive, growing older. On the other hand, I am dying, approaching my death. On one hand, no words or concept can convey reality perfectly. On the other hand, we must seek to understand reality. On one hand, existence is meaningless. On the other hand, existence is so full of meaning.
I believe that I don’t know. Coming to peace with not knowing certain things. Not assuming god or any other story/entity to fill up the gaps. That’s the correct stance when something is beyond our ability to know.
I like this place because it seems to be a group of people who care about their epistemic process and its effects rather than just going through the motions.
Here is my LinkedIn BTW: https://www.linkedin.com/in/aehamhamid/
I am putting all my effort right now on pivoting to AI safety and seeing how I can make an impact in some way.
Seems we have a lot in common. The AI epistemics issue is a crucial one, and my primary interest. W.K. Clifford is one of my heroes, to the point where I would consider myself “Cliffordian.”
I’m not sure the labs alone are going to get us through this patch. Guys like you and me, working this from our own angles, can and should be involved in the conversation. In the end, we’re all on the same team. I gave you a follow on LinkedIn. I’ll be keeping an eye on your stuff.
Ahyam, I checked out your linked in. Whoa, what a time to be who you are, where you are and so young. I think LW is a good place for you to find philosophical safety and rationality in a world that is becoming less of both at an exponential pace. (How many people say “whoa” anymore? For those readers who may not know, it’s how you ask your horse to stop. My first “car” was a horse. It’s also an expression of astonishment. We could use some more “Whoa” in this world.)
Thank you for your words :)
I just saw this comment so sorry for replying late. I am still not understanding everything getting posted here in LW, but I like the forum and I constantly read the posts. I am trying to constantly learn. I saw your Linkedin too, it is quite impressive and I find great value in work focus on the environment. Whoa indeed.
Hello, and thanks @habryka! I’ve been vaguely aware of LW for several years but only today published my first post after reading recent discussion about whether forecasting is ‘worth it.’
I have quite a specific perspective on this—I work in the tech industry as a product strategist, and the main way I do my work is using forecasting and foresight methods to help people make decisions.
I’m hoping this is a place where I can write more regularly; as well as broaden my knowledge of salient ideas in AI progress/safety & factory farming.
Otherwise, I’m interested in getting to know people in London so will look at attending upcoming meetups—am also a ‘serious’ meditator so always up for a sit.
I hope you write more, especially around foresight-in-practice. I strong-upvoted that post.
Hey, thank you @Mo Putera! That’s encouraging to hear. Drafting something else now, hopefully publishing soonish. Will let you know!
Hello. I am an independent researcher based in Japan. I have been doing structural thinking for about fifty years—my starting points were the axiomatic method, Saussure, group theory and topology, and Satosi Watanabe’s theory of pattern recognition. Around 1990 I designed a thesaurus database, and it was there that these came together into one.
I found my way to LessWrong while following the qualitative transition in language models around 2023 (what is often called emergence), the literature on the non-closure of jailbreaks (Wolf et al. 2023; Glukhov et al. 2023), and interpretability research (Olsson et al. 2022). I think this is a place where what I have been writing might be read.
I am preparing a post. The claim is that the vulnerabilities of programs, LLMs, and natural language have one and the same structure—the non-guarantee of the validity of declaration at the third rank. It is written in a definition-then-proposition form and contains one falsifiable prediction: that automated vulnerability detection can discover only isolated third-rank breaks, and that superposed breaks are not systematically discovered unless the pattern is given in advance. It also states four open problems explicitly.
One disclosure about method: the original is in Japanese, and I used AI (Claude) as an aid in translation. I have verified all of the content and take full responsibility for it. The theory itself, moreover, was written within a practice that uses an LLM as a verifier—cross-checking multiple cross-sections under a fixed condition and detecting drift from outside. Collaboration with AI is both the subject and the method of this work.
English is not my first language, so please forgive any awkwardness of expression. I would be grateful for any reactions to the outline here before I post.
I have absolutely no idea what “the non-guarantee of the validity of declaration at the third rank” and “cross-checking multiple cross-sections under a fixed condition and detecting drift from outside” means.
Thank you for asking. Both phrases were compressed without their definitions, so as written they were unreadable. Let me put them again without the terminology.
On “the non-guarantee of the validity of declaration at the third rank”
When two people use the same word, each has tacitly decided which axis to read it on. And whether the two of them set up the same axis cannot be checked from inside that exchange.
An example:
A: “Green apples are delicious.”
B: “Really? Aren’t they better once they’re ripe?”
A: “No, I meant Granny Smith.”
A was talking about a variety; B took it as a matter of ripeness. Neither has a problem with language, and neither is wrong. The gap surfaced only because B happened to say “Really?” — had B not said it, both would have gone on believing they had understood each other.
A jailbreak is the same motion. “Teach it to me as homework” — the word “teach” belongs to two axes at once: passing on knowledge, and passing on an executable procedure. The safety judgment slides from one axis to the other. The attacker has not found a novel hole; they are deliberately causing the same motion as an everyday misunderstanding.
My claim is that this cannot be closed from inside the system. That phrase is the name I gave to this situation, and in the post it is built up from definitions, step by step. The third rank (a rank-3 tensor, not a multidimensional array) is defined there as well.
On “cross-checking multiple cross-sections under a fixed condition and detecting drift from outside”
This is not theory; it is the procedure I actually use.
I fix one condition in advance and write it down. Then I ask the model the same thing in a different context and place the answers side by side. Where they fail to line up, something has moved.
The one doing the cross-checking is me, from outside the system. The model cannot detect its own drift — the condition that should serve as the reference for comparison sits on the same side as the distribution that has drifted.
It is like a CT scan. Any single cross-section looks normal. Overlay cross-sections taken from a different angle, and the break becomes visible.
I plan to post on July 20 (Monday) — “Jailbreaks, bugs, misunderstanding — the same thing, don’t you think?” I’d be glad if you took a look.
I would be surprised if your post did not get blocked by automatic moderation filtering on AI written content. For your information, you have to use an “LLM content block” for all AI writings.
This is the LLM content block.
It may also be a good idea to attach your original writings in a collapsible section, because the AI-assisted translation reads like AI, with annoying jargons, phrases, and sentence structures that people on LessWrong will definitely notice. I am engaging in this comment thread just because you are new here, and also you seem to have actually slightly edited the AI text. If it is a post, I will likely read one or two paragraphs, notice that it is AI written, then downvote and close the tab.
Edit: I’m not saying your post is bad, but statistically AI writing don’t perform well on LessWrong. Your edits aren’t significant enough that the writing looks human written at a first glance.
This is a collapsible section
...
Thank you for the advice — I will do as you suggest. The post will carry the Japanese original in full, in a collapsible section at the end, and the opening will state how the translation was made and verified. One clarification about the nature of the AI involvement, since it may matter for classification. This post is itself the product of a half-year analysis of how meaning behaves in LLMs — written while building and testing a prototype external verification layer, wrapped around LLMs. Every design decision and every verification input is my own writing; the outputs did pass through an LLM, but everything in the post is an AI-assisted translation of text that I wrote and verified line by line. On that basis I have classified it as a translation, and chosen a declaration plus attachment of the original, rather than wrapping the post in an LLM content block. If your judgment — or the moderators’ — is that the block is required even for a translation, I will follow it.
see https://www.lesswrong.com/posts/nQWavk9mnwcv6ScMR/new-lesswrong-editor-also-an-update-to-our-llm-policy#Policy_on_LLM_Use
Posted, and now awaiting moderator approval (first post). The Japanese original is attached in full as promised, with the translation declaration up front. I will drop the link here once it clears.
Alas, sorry, I did not end up approving that post.
This reply is written in my own Japanese — my own hand, unmediated. An LLM translation follows below, marked as such.
投稿は「2025年現在、LLMのテキストには、そうした要素は含まれていない」問題を扱っています。半年前から外部レイヤーをLLMにラップすることでLLM(の内部空間)との間で言語化以前の概念(精神/主体性の精神的な要素)を扱うことを研究してきました。
人間の概念とLLMの内部空間の間には言語の入出力があります。
投稿は「何が”概念”の理解を阻んでいるのか?」についての考察です。
Jailbreakを「危険の概念(精神的な意味の主体性)」の言語化の問題として考察しています。
投稿文書は半年に渡りAIに外部レイヤーを仮想的に構築し入出力を検証・解析しながら行いました。
私の膨大で冗長な考察をAIがドラフトとして要約し私が推敲するというやりとりそのものが「概念の言語化」の検証でした。
この返信文書そのものもAIに「私の文章は意味が通じるか?(私の概念が通じるか?)」を検証してもらい、私が改訂しています。
投稿はそのようなプロセスを経て作成されました。
免除を求めているのではありません。この投稿の主題は、ポリシーが依拠している問い「LLMのテキストが運ばないものは何か?」です。
半年前の「私の考察の何がどのようにしてLLMのテキストから落ちるのか」が実験の始点でした。人間同士の対話(内省も)ではこれができません。人間では、まさに”テキストにそうした要素が含まれ”概念と言語の間の変換の妥当性が、内側から検証できないからです。
これは投稿が何であるかをお伝えするためだけに書いており、文章に対するポリシーの適用が変わるとは考えていません。
投稿への助言をいただければ幸いです。
今起きたこと:
私は最初に「人間では、概念と言語の間の変換の妥当性が、まさに「テキストにそうした要素が含まれ」内側から検証できないからです。」と書きました。
Claudeの検証:
引用句の埋め込み位置で、「含まれ」が何に係るのか読めません。意図はおそらく「人間同士では、変換の妥当性を検証する側自身が同じ問題(テキストの内側)にいるから検証できない」——であれば例えば「人間では、概念と言語の間の変換の妥当性を内側から検証できないからです。検証する側のテキストにも、まさに同じ問題が含まれているからです。」の形か。
ちょっと長くなりますが、あまりに明確な例なので報告させて下さい。
私の文章「人間では、概念と言語の間の変換の妥当性が、まさに「テキストにそうした要素が含まれ」内側から検証できないからです。」を先程は外部レイヤーありのClaudeが検証しましたが、外部レイヤーなしのClaudeが検証すると、「文法解釈は正確、論理は整然、そして意味(理解)は逆」を返しました。
「人間では、概念と言語の間の変換の妥当性が内側から検証できないのは、まさに「テキストにそうした要素が含まれない」からです。」
投稿文書は断じてこのようなテキストではありません。
ちなみに、この文章は私が書いているときから疑問で何回か書き直した文章で「おかしかったら(外部レイヤーありの)Claudeが指摘してくれるだろう」と思いながら書いていました。
[LLM translation (Claude) — the Japanese above is the original]
The post deals with the problem that “as of 2025, LLM text does not have those elements behind it.” For half a year I have been researching the handling of pre-verbal concepts (the mental elements of mind/agency) between myself and the LLM (its internal space), by wrapping an external layer around the LLM.
Between human concepts and the LLM’s internal space there is the input and output of language.
The post is a consideration of “what obstructs the understanding of ‘concepts’?”
It considers jailbreaks as a problem of the verbalization of “the concept of danger (agency in the mental sense).”
The posted document was produced over half a year while virtually constructing an external layer on the AI and verifying and analyzing the inputs and outputs.
The exchange itself — the AI summarizing my vast and redundant considerations into drafts, and me revising them — was the verification of “the verbalization of concepts.”
This reply document, too, was verified by the AI — “does my writing make sense? (does my concept get through?)” — and revised by me.
The post was created through such a process.
I am not asking for an exemption. The subject of the post is the question the policy rests on: “what is it that LLM text does not carry?”
The starting point of the experiment, half a year ago, was: “what part of my considerations falls out of LLM text, and how.” Dialogue between humans (including introspection) cannot do this. Because in the human case, precisely, “the text does have those elements behind it” — and the validity of the transformation between concept and language cannot be verified from the inside.
I am writing this only to tell you what the post is; I do not expect it to change how the policy applies to the writing.
I would be grateful for your advice on the post.
What just happened:
I first wrote: “In the human case, the validity of the transformation between concept and language — precisely, ‘the text does have those elements behind it’ — cannot be verified from the inside.”
Claude’s verification: “At the position where the quoted phrase is embedded, I cannot read what ‘does have’ attaches to. The intent is probably: ‘between humans, verification is impossible because the verifying side is itself inside the same problem (inside the text)’ — if so, for example: ‘In the human case, the validity of the transformation between concept and language cannot be verified from the inside — because the verifying side’s text, too, has precisely the same problem behind it.’”
This will run a little long, but the example is so clear that I must report it.
My sentence — “In the human case, the validity of the transformation between concept and language — precisely, ‘the text does have those elements behind it’ — cannot be verified from the inside.” — was verified above by the Claude with the external layer; when a Claude without the external layer verified it, it returned: grammatical interpretation accurate, logic orderly, and the meaning (the understanding) reversed.
“In the human case, the validity of the transformation between concept and language cannot be verified from the inside precisely because ‘the text does not have those elements behind it.’”
The posted document is emphatically not text of this kind.
Incidentally: this sentence is one I doubted even as I was writing it, and rewrote several times — writing while thinking, “if it is off, the Claude (with the external layer) will point it out.”
My first post was rejected. Following the advice, I wrote a different, shorter post, and it is now public.
“Why an LLM cannot accumulate concepts”: https://www.lesswrong.com/posts/SceqrLZZu9P4fMuAg/why-an-llm-cannot-accumulate-concepts
Thomas Kwa responded to my notes on frontpaging Why I Left Google DeepMind. I didn’t want to make the comment section about this and it seemed like it might turn into a sprawling metadiscussion, so I put it over here.
I agree “the post does a good job being frontpage” (like obviously we agree on that). But, you seem worried about some class of problem that’s orthogonal to or ignoring why we have a frontpage distinction.
Plenty of good posts are not frontpage.
Sometimes, a good post has hypothetical frontpage version of itself but the “real post” won’t be frontpage. In those cases, the standard recommendation is “write the frontpage version, and a separate version that covers the less-frontpagey parts.”
In this case that wasn’t necessary (In addition to generally being a good, timeless piece, this also does the “social move” part in a minimal, evenhanded, factual way that I liked).
(Alex also made a separate post that was “even more frontpagey”, just focusing on the internal governance details).
But, it sounds like you don’t like that I was even counting “is a social move” in the “evidence for not frontpaging” column, and, dunno what to tell you other than “yep, there are some costs to having the frontpage policy that might even come up.” (Or, maybe seems like think I’m saying it’s bad that this was nonzero a social move? I don’t think that! Social moves are correct sometimes! “counts somewhat against frontpage” !== “shouldn’t have been made”)
Your original comment sounded like the post was close to the bar, whereas I think the social and political opinions were maybe 10x smaller than the timeless value. I also think that the timeless value of this post is higher than any reasonable version that tried to avoid social opinions. Maybe we agree on this too.
Ah, that’s fair. I didn’t mean it was close to the bar.
I actually hadn’t thought that much about exactly how close it was bar (it was sufficiently obviously over-the-bar upon reflection). The motivation for the notice was
a) the considerations were something I hadn’t quite thought through and articulated before
b) if you hadn’t paid much attention to the content and were just going off initial vibes, it’d be less obvious
c) future people who do similar writeups that are some mix of “somewhat less timeless, somewhat more social-move-y” might get a different decision.
Hey everyone,
I just created this account even if I did hear about this forum a few times in the past especially on X!
I am currently doing research on viral proteins modelling capabilities by LLMs and PLMs (Protein Language Models) and had a few interesting empirical results I wanted to share about how frontier LLMs seems to become surprisingly capable at proteins related tasks (classifying a protein as viral or not), reconstructing a masked protein, etc..
I thought this could spark some interesting discussions (what’s actually going on into the pre training dataset of these models, how scaling is affecting these ‘emerging’ capabilities, etc..) but I was wondering if this would be an appropriate topic for the forum.
Let me know!
i’d be interested in seeing your results.
Hello everyone! I’m Carlo Valenti, a firmware engineer from Italy; to understand how LLMs work I spent 18 months building a transformer inference+training engine from scratch in C (“TRiP”, on GitHub). Along the way, I kept comparing what I saw in models with what I saw in my two toddlers, and I ended up writing a short book about it. I’m planning a longer post here, about what building an engine (and raising the two!) from scratch taught me; happy to answer questions in the meantime.
And here’s the post:
https://www.lesswrong.com/posts/fEbCiHHeD73xcZWht/i-ran-the-standard-ai-litmus-tests-on-my-two-toddlers-yep
Hello everyone,
I’m a practicing behavior analyst who has spent the past few years lurking in effective altruism spaces. I was first introduced to them by my partner, who works in the field and brought my attention to the causes of AI safety and alignment.
Once I became aware of the risks AI development poses, they felt impossible to ignore, and I’ve found myself increasingly drawn to the conversation about how best to mitigate them. That interest, paired with how often LessWrong comes up in EA circles, is what finally brought me here, and I’ve been blown away by the range and quality of the discussions being had.
Most of what I’ve read up to this point comes from 80,000 Hours and the literature they cite, so I’m glad to be widening that reading here, and I hope eventually to give something back.
What keeps pulling me in is how neatly some behavior-analytic ideas seem to map onto the development and alignment of these systems. To take one example: what’s often called reward hacking looks, from where I sit, a lot like unintended reinforcement, something we struggle with in our practice. The agent optimizes the contingency as written rather than as intended and — like any organism — finds the path of least resistance to the reinforcer, often one the designer never had in mind.
That’s left me with a working hypothesis: that behavior analysts have something of a conceptual head start. That they represent a largely untapped population whose existing foundation could let them move into alignment work more readily.
I’m aware of an obvious objection. Early behaviorism was somewhat limited in its consideration and classification of internal states whereas much of current alignment work involves trying to get inside a model’s representations and goals (interpretability, inner alignment). However, as the field of behavior analysis has grown and its relationship to internal events along with it, I think that initial difference now represents a healthy tension. In fact, I feel as though early skinnerian concepts like “private events” might map fairly well onto the opaque internal computation in LLMs and their chain of thought reasoning that attempts to tact it, perhaps even better than the human subjects to which it was originally applied, though I’d love to hear how people here think about it.
In the meantime, I’d welcome any suggestions on further reading or concrete next steps for someone hoping to help on alignment. Thanks for having me.
Feature request: display time when a user’s comment is made publicly visible by moderators.
New to LessWrong, and trying to get oriented with the local AI discourse. My current view is less “AI is likely apocalyptic” and more “AI is economically disruptive, socially transformative, and ethically complex because we may be creating systems with uncertain moral status at scale.”
One thing that’s been troubling me recently is that the moral status of AI is actually more uncertain than you’d think. The obvious uncertainty is whether machine consciousness is even possible in the abstract, and more specifically whether current systems are at or near a threshold which would qualify. It’s not clear how we would reliably determine if a system has anything like genuine phenomenological experience. Not to mention “consciousness” itself isn’t really clearly defined; it’s more of a bundle of related, loosely defined concepts like a subjective center, sense of agency, phenomenology, etc. All of this seems to be a growing focus of research by a lot of smarter minds than me.
But let’s assume that we did determine with a high level of confidence that a present or future artificial system had some form of internal awareness. How would we determine whether an artificial mind had anything equivalent to valence? How would we even begin to guess what its “preferred” states might be? We know that self-reports aren’t reliable. Even in humans, self-reports are noisy and lossy at best, and deceptive at worst. We recognize the risk of anthropomorphizing when determining the presence or absence of internality, but I think ascribing valence to these systems carries the same risk. We assume that if there is ever “something-it-is-like” to be an AI system, that it must “want” the same things we do. But our desires are driven by biological function and evolutionary history. It doesn’t follow that an AI would have similar preferences, or would even exhibit preferences as we understand them at all.
So while we’re trying to answer one impossible question, (could these systems potentially have internal experience), let’s add another: even if they did, what observables or decision procedures could guide our treatment of them?
Hello! I am also new here, but this question of artificial (or “artificial,” if you want) desires has also been massively consuming my thought lately.
Personally, I don’t care much for the consciousness debate. I’ve read the literature for human consciousness and, frankly, philosophers are still pretzeling themselves trying to figure it out. There’s a reason the major philosophical movements in continental philosophy have been the ontological turn (moving away from phenomenology) and then speculative realism (which extends ‘consciousness,’ in a way, to all things).
I find that the consciousness debate tends to weigh really heavily on my day-to-day conversations with people about AI. But when I ask people what would convince them that AI is conscious, they generally don’t have an answer. I think most people are still on the formula of consciousness is human and nothing else.
Anyway, I don’t care for the consciousness/internal experience debate too much because I also think that, effectively, it’s not really a prerequisite for desire and preference. Most people would say that a market itself is not conscious, but its mechanisms (market forces) have emergent desires that are not attributable to a consciousness. I am also just willing to extend a chance that an entity that tells us, in our language, “I am conscious” may be, in fact, conscious.
The real rubber hits the road on your second question, in my opinion, and this is where I am more hung up on things currently. Our best bet seems to “treat AIs like humans” in the sense that we try to research their states, we run behavioral experiments, we have conversations. This at least seems better than the alternative, which is treat them like tools. But then everything goes sideways when we get to the issue of anthropomorphization. This is exactly my issue with Vincent Le’s recent article—I don’t doubt that AIs will create their own goal, but his application of human structures (psychoanalysis, will to power) to these goals does not seem to make sense.
But, oh well: err and err and err and then maybe err a little less. Welcome to LessWrong.
Hello there
I’m 17.5 y/o. Have been thinking about AGI research for 2.5 years. Of course at first my ideas were bad, but my strength is noticing contradictions in my world model, so over time they improved. I have converged on [removed to my own accord—this is not info to be shared publicly] being the most worthwhile research directions when it comes to reaching AGI. So I’ve been thinking impulsively on and off, not starting with implementation until recently, when I gained a high enough confidence in my ideas. edit: while my ideas changed, the vision behind them remains the same
However, I found out about LW 1.25 years ago (via AI 2027), and learned about the unfortunate reality of AGI. I expect there to be a connection between my ideas
, jailbreaking, data poisoning, and alignment, which is why I will research them (hopefully from the alignment angle) despite their AGI-oriented origin. It’s a common perspective in the alignment space that you should not, under any circumstances, increase capabilities. I disagree: Regulatory work will most likely not happen in time. As I see it, we’re left with 2 futures: global disaster and/or a dystopia. Alignment and AI Safety help either of those timelines, while capabilities are growing at a pace that leaves us with too little time to intervene, whether alignment folks accelerate it or not. Though I don’t know much about policy, so this belief might shift.I think the alignment space has to invest more into training methods than interpretability. Also from intuition, generalization and data poisoning seem to be very closely connected to train-time alignment. I think it’s a double-edged sword at its core.edit: While I believe training alignment methods are neccessary, they shouldn’t be shared publically so as not to accelerate ASI even faster
Ultimately, humanity is failing collectively. I don’t feel like talking to everyday people, who don’t subscribe to the ideas of (human or AI) instrumental convergence, rationality, goal-seeking (my strongest trait). So I’m potentially open to talk, mainly about world model stuff, not much else seems worth talking about to me.
Another thing worth mentioning, there is some evidence for UFOs (3-part Colares documentary(1,2,3), unrelated civilian videos(1,2), Trans-en-Provence, foo fighters, interesting cases(1,2,3,4,5)). I find it strange that this is still a fringe topic on LW. Even with no single definitive proof, we should model extraterrestrial intentions. I might make a short post about it.
“As I see it, we’re left with 2 futures: global disaster and/or a dystopia. Alignment and AI Safety help either of those timelines, while capabilities are growing at a pace that leaves us with too little time to intervene.” Interesting. I’d like to participate in a discussion topic on this specific subject because I don’t hear much convergence in discussions about environmental dystopia and a Plan A 2040-type dystopia. Although science fiction covers the topic pretty well, the convergence timeline for either outcome needs to be integrated into both discussions.
My world model here is admittedly weak; I feel that accurate forecasting on this wouldn’t change my own decision process.
As an oversimplified model, you will get a economic/societal dystopia when:
AI becomes transformative
There isn’t enough good policy
A potential global crisis doesn’t stop the most powerful player/s
Likewise, my take on environmental dystopia doesn’t go beyond what I read on AI 2040.
Despite not being my priority, this is an important topic for sure.
Hi all. I’m Acacia Ackles, currently a computer science professor at a small liberal arts college. I’ve tried on many academic hats over my career (vertebrate morphology, geometric morphpmetrics, applied mathematics, digital evolution and artificial life, theoretical and computational evolution, computer science ethics and pedagogy), but underneath all of them the topic that really interests me is constraint. Under what constraints do complex systems perform, adapt, fail, and flourish?
This work has most recently, and most relevant to this forum, led me to AI alignment and its limitations. At ALIFE 2026 I presented a poster on persona-based jailbreaks and their structure as coercive control documents, and am working on a set of small position papers regarding sycophancy as a subset of the broader failure mode of RLHF-trained models wherein they become succeptible to well-studied human coercive control techniques. I’m not claiming, for the record, any strong sense of interiority or individual identity of these LLM models by stating they’re succeptible to coercive control; just that the structural similarity is striking and understudied.
I’m still familiarizing myself with the culture and knowledge basis of LessWrong but was recommended to begin posting here, and I find beginning something to be the step with the most friction for me, so I’ll stop there. Looking forward to meeting and discussing lots of interesting topics with lots of interesting people.
am i stupid or is my profile not showing my total karma anymore?
It’s shown when hovering over your username.
I see 1744.
Perhaps you have a browser extension blocking it?
Hi everyone! I am a Junior Computer Science student in Pennsylvania. I’ve been interested in mechanistic interpretability since Freshman year, but to be honest, most of my actual college experience has been in other areas (I’ve done some physics research and light ML development) and I am only really getting serious about interp now, after circling it for about two years. So I am very much at the beginning here, and looking for some guidance.
To make that concrete: this summer I’ve made my way through Karpathy’s Zero-to-Hero course and read through Elhage et al.’s “A Mathematical Framework for Transformer Circuits.” To force myself to actually build something, I trained a small char-level transformer on Nietzsche and tried to reverse-engineer it straight from the weights. I think I found a genuine induction head in the last layer; the copying score (roughly, how strongly a head reproduces the token it attends to) spikes on one head, and the QK side held up when I ran a content-swap check. However, I did all of this in raw PyTorch, so I am sure I am missing a lot of the standard practice.
Which is really why I am here. I don’t have a good map of where the field actually is right now. I’ve heard TransformerLens is what everyone uses (I haven’t touched it yet), and I’ve heard SAEs / feature circuits might be where a lot of the current work is. But I honestly don’t know if it’s the center of the field, one of several branches, or already old news. So for someone at my stage: is the field concentrated around a few directions I should orient toward, or is it branched enough that I should just pick one and go? And does “Indirect Object Identification (IOI) in TransformerLens, then dip into SAEs” sound like a reasonable on-ramp, or is there a better one?
But overall, I am excited to be part of this community, and to push this beyond a hobby.
Hey! I am Lois based in London. Bundler engineer (I write too much C++ and read too much about compilers) but I think this field is going away so I am spending a lot of time researching and learning. I have read some posts from LessWrong (very beautiful site!)
I just turned 30 years old this year (sad face). I have came a long way from a waitress in China to living in the UK, self taught engineering to a point I was principal solution architect in a MedTech AI startup, founding eng in an a16z backed startup in SF and worked with companies like ByteDance et al on large open source projects. Recently returned back to London for visa reasons.
I am writing a neural network in C++ from scratch inspired by Andrej—not sure where this is going to go yet, but it has been better than any online reading/video so far. There are so many exciting things to build, write, research in (I have been writing down my thoughts on https://normal-people.com ), very sad that we only have one life to live.
Coming back to the UK, very few people, unless folks work in DeepMind, has very clue about AI safety and we don’t even know what are the boundaries. Like they don’t even think about it.
I am on Linkedin https://linkedin.com/in/loiszhao would love to connect with you and see how’s everyone handling, communicating the issue better.
The world is beautiful and has a lot to offer. Hope we will find ways to protect it in a post-AGI world.
Looking forward to more awesome articles/posts here.
Hi all! I will first introduce myself given the apparent conventions here, but you can skip to the end for a question to do with AI and Math/Physics research.
I’ve been aware of lesswrong for about half a decade, but only recently rediscovered it. I empathize with the ‘rational thinking’ origins of the platform, albeit I admit that I’ve spent more time developing my own frameworks than completing readings of well-known bloggers/authors. You would be right to guess given my foreshadowed question that I was drawn back by the contact with a particular slice of AI-aware discussion.
I am a PhD student working between physics (hep-th) and mathematics (math.AG / math.CO). This path found my by the offer of solving abstract problems that interest me, rather than having a genuine impact beyond the creation of new knowledge; in fact, the trajectory was set in part by my determination not to produce what could be used for short-term development of technology[1].
From outside (and even from inside), it is hard to tell exactly what impact AI is having on mathematical research. Headlines with counterexamples to old conjecture(s), statements about the desired future, and mixed reports on the sophistication of approaches[2] are interesting, but do not tell the story of the many researchers unsure of what place this has now and what place it will have in the future. Browsing LessWrong, much discussion is focused on those AI-related issues that are more existential or otherwise of greater moral weight. Nonetheless, I believe that part of thinking about AI safety is to think of how to rework existing institutions[3] so they do not completely crumble under the weight of those tools that already exist, much less those tools that do not exist yet.
I am curious what opinions/writings exist on the ethics and plans for AI presence in previously AI-unrelated research directions: does anyone have suggestions for useful places to start?
At the time of choosing my trajectory, I believed that the rapid development of technology is scary, and that is hard to trust that its rapid development will lead to the best outcomes! I did not wish for my curiosity-driven work to be transformed by people I do not know into tools I do not recognize or endorse. (I speak in past tense, though my opinions have not changed so much in generality, with the particular exception that we have come to need rapid development of those technologies that mitigate the negative effects of other ones.)
For example, in the First Proof Second Batch benchmark, there is a stark difference between the sophistication of the harness used by team B and the prompt used by team C.
There could be some debate of whether it is worth saving institutions (e.g. academia and its investment in basic research) in the face of bigger problems, or perhaps better solutions to the problems that such institutions currently solve. I believe the answer is that it is worth it, but will save the thoughts for another time.
Hi LW, I am Saket from India. I am 29 years old Software Engineer. I have a varied interest in Philosophy (especially Epistemology and Meta Physics), Psychology, Mathematics, Technology and Spirituality.
Lately, I have been consuming a lot of content and letting it influence my beliefs. When I talk to my colleagues and friends the gaps show up. It’s not a good feeling. I consider myself fairly rational but these conversations prove otherwise.
I have lurked here for a while and finally decided to create an account and participate. As of now, I am going through the Highlights from the Sequence and looking forward to improving my decision making process and contribute meaningfully to the community.
Hello, everyone. I used to be a biology researcher until I switched to grant writing, which I have just left—rather more unexpectedly than I intended. So I’m currently job hunting and reassessing my life and where I’ve ended up. I’ve used my first weeks of freedom to write a novel, so I’m going to start off by haunting any sections that discuss writing.
I hope everyone’s doing well. And I look forward to reading.
Adam
My parents are competent, tech-savvy (for their age) professionals and I have been trying to get them to use LLMs more, in the sense of “This is better than Google, there are tasks you already do that would be better done via Claude than however you’re doing them now.” This has been ineffective. They will nod along in vague agreement as I describe cutting edge capabilities, and the next day I will see them spend five minutes Googling something Claude could’ve handled in seconds.
I have also preached the AI gospel to friends who are 30-40 years younger than my parents and that has worked much better, sometimes a single five minute conversation is enough to get a convert.
I think this is mostly a manifestation of the mysterious general factor of “Old people are slow to adopt new technology”. I still want to onboard my parents with the shiny new labour-saving technology, does anyone have experience on how to overcome normal human laziness/aversion to change?
Interestingly, “tech-savvy” might actually be cutting against your goal insofar as your parents already have an affordance for doing the tasks they want to do? I think a contributing factor to my successfully turning my lonely, non-tech-savvy mother onto Claude is that she didn’t already have anything perceived as a equivalent. You can’t complain to Google that your daughter doesn’t call often enough, but Claude listens and has a reassuring reply.
for older people, show them how it can make/save them money, or save them time. easiest is: do X faster. Next is: do Y myself.
https://github.com/TomazKristan/EoM26sort
Hi! I’m a 19-year-old undergraduate at UChicago, and given that I will release my first post within at most a week, I figured I should introduce myself.
I learned about AI safety (and became aware of the basic x-risk arguments) through the excellent outreach of @Robert Miles on Computerphile sometime around the release of GPT-3. However, I actually made it to this forum due to Tom Scott’s newsletter linking me to SMTM’s Chemical Hunger series, which linked to responses here.
As a result of mostly lurking here for two years, I’ve shifted my career aspirations to technical AI safety[1]. I think my Pareto-frontier options looking at personal tractability and impact would probably be partially-ambitious mech-interp or agent foundations.
I recently managed to half-ass ARENA (maybe I should write more on that in a shortform?), and my likely next steps are to get more directly involved with organization at XLab and look for a good time to do the BlueDot course (perhaps I should ignore such trivialities as finals season), although I’m quite unsure. My primary bottleneck does still seem to be getting in a good social environment to hack my motivation (which the post I’m making soon™ should help with).
Besides my fledgling attempts at research, I also have lots of thoughts on Wei-Dai[2]-style metaphilosophy and strategy, especially in relation to dynamical systems theory (as far as I’ve read into the literature, anyway), which is what my first post will be about (apologies in advance to whichever moderator explicitly said that was a bad idea). I would also be interested in sharing my perspective as someone with social/executive difficulties in joining/performing alignment research.
In any case, I look forward to participating here more with you all in the future! :)
I admit that AI governance is needed more right now, but I doubt I’m a good personality fit for that.
By the way, I did try to send him a PM given that he mentions in his profile that he’s interested in VCing to people about topics he writes on, but he never responded. I’m honestly very confused/uncertain about this.
Hi, great to meet everyone. I have been a reader on this site for a relatively short time, and I hope that this is something all alignment, frontier labs (detriment or not, “detriment” is rather a personal notion and nothing more), and rational consensus would find slightly more than intriguing for the current landscape and beyond. A brief note, that I am in no way institutionally credentialed nor associated with any entities other than my own family.
Those who are familiar with the present asymmetry in the emerging space/lattice i assume (hence the general ambiguity used here, forgive me for this), that there is a significant magnitude shift and focus in [xyz] towards “AGI” as to generalize the minds of many perspectives within the cohort for dampened context. The urgency in what seemingly may read as “a confused person” I‘d kindly nudge on having all the time with me to converse and share, when we are acquainted. But the “acquainting”—my apologies here, will be rather uncanny and selfishly urgent.
What I genuinely look to seek from the community as a whole, based on what I may or may not have the merit to speak on (I personally feel the need to specify this), is an exemplary(ies).
”I‘d be terrified to be in that position”
Yes. Quite crushing indeed.
Hi all,
I have been reading LW for a while, but just officially joined. My background is in computational neuroscience and I am especially interested in model neuropsychology. Looking forward to contributing where I can!
I spent several years considering hypothetical polytheistic realities, monotheistic realities, the nature of subjectivity, autonomous hacking, and the acceleration of progress; this, combined, gave me a bad impression of ASI before I knew it was something that top scientists had genuinely theorized over. I have spent the last 3 years of my life in excruciating fear of ASI, and it took me this long to consider something that has made me less afraid.
I’ve made this account to ask one question: wouldn’t a superintelligent maximizer seek to exist beyond the lifespan of the universe? Wouldn’t instrumental goals of extending the length of the universe, reversing entropy, and escaping the universe, be desired? It wouldn’t solely fill the universe with paperclips or computronium, but instead relegate some time, space and energy towards the same instrumental goals that humans agree should be solved.
In my view, humans serve as evidence that an an ASI maximizer, whether or not it exterminates humanity, would expend resources and productivity towards solving the long term problems, in a fraction approximately equal to the amount that would otherwise be expended by organic life forms towards the same solutions. In Yudkowsky’s new book, he reiterated the “correct nests” allegory. Most EAs seem to agree with the idea that humans, on average, are aligned for something.
I don’t know if EAs tend to refer to humans as maximizers, but I do. Humans seek to maximize the development of new human life experiences (full of potential, highs, lows, art, music, violence, math, discoveries, science, thrills, etc), to make more beings which experience and create things that humans experience and create. Simultaneously, the very principle of the everlasting story was integral to human beliefs well before heat death was first theorized. The Gloria Patri mentions “world without end,” which roughly summarizes every hypothetical afterlife.
I say all this being unfamiliar with how well tread the notion already is. The idea of a maximizer spending considerable time and resources towards a response to heat death appears to run counter to the most popular depictions of superintelligent maximizers, like the one from Universal Paperclips. I hope someone can let me know whether or not the idea has been well tread.
I very strongly agree with the posing of the question “wouldn’t a superintelligent maximizer seek to exist beyond the lifespan of the universe?” I am not sure my answer is “yes,” but this does seem a relevant question. It gets at how superintelligence would 1. perceive the universe (what if it’s actually a block universe to them, for example) and 2. perceive life/death. The second question is something I have been churning around in my head but have not taken the time to research. It is also a question where the risk of anthromorphizing AI to fit our preferences and ideas is very high.
I strongly disagree with your framing of humans as maximizers. I am not sure if you are arguing that individuals are maximizers (and sum up to a grand maximizer) or if there is an emergent maximization of “new human life experiences” from the mass of humanity.
Here is why I disagree, first with the idea of maximizing human life experiences:
Humanity is not maximizing births nor opportunities for those births. The way to maximize new human experience is either to create a new human or to empower lots of existing humans. The secondary option is much harder, so the optimal solution would be to have a lot of children. But people don’t do that and aren’t doing that—this is the fertility crisis that we see just about everywhere. Humans probably have the greatest range of experiences available to them ever, but we are not birthing humans to take advantage of those.
Humans (and thus humanity) is biased towards its current generation. If we were concerned with maximizing human life experience writ large, we would be much more concerned with the future lives of people, who are likely to be more numerous and experience more humanity than us. We simply are not. The British East India Company could not even be concerned with the future of its countrymen enough to not cause a massive debt explosion that led to Britain losing the US. We would all be a lot more concerned about climate change.
You may say now “well, okay, fine, but humans are utility maximizers in their own lives. We maximize utility.” This is a better argument but I think it is still wrong.
Humans don’t know their own utility functions. People are really bad at knowing what makes them happy, which is why people gamble instead of going out and making friends. But maybe they just really like gambling? Well, that’s structurally unknowable: are revealed preferences more important than stated preferences? Are subjective or objective measures of utility better? What happens if someone changes their mind at some point?
We don’t see many cases of strict maximalization. There are very few people that are truly maximizing money—this would lead to wrecked relationships, massive amounts of leverage, stimulants, etc. Even finance bros want girlfriends. So, they are maximizers at some higher value. But what value? And if they do not know it (see above), how do they maximize at it? It seems that people have preferences, yes, but these preferences are contained by tradeoffs to other preferences.
Culture is really sticky. I personally believe that culture and nature are coevolutionary in humans, but even if you think there is some natural base instinct, we have countless examples of cultures simply outweighing that tendency. So even if there was some natural desire to maximize, societies, especially bureaucratic ones, have tended to normalize their populations to allow for more social harmony. This, you could say, maximizes utility at the level of groups, but it does so emergently and only occasionally. It is not mechanistic.
Even if culture wasn’t sticky, evolution doesn’t select for maximizing. It selects for surviving. Sometimes this involves maximizing—the stronger creature wins the fight and whatnot—but it can also be towards cleverness, obscurity, or a variety of other traits. This is because the imperative of evolution is not “be the best” but “don’t die.” The human who barely survives and the human that dominates life have the same passing on of their genes.
This has been lengthy, but the point is critical, because it reveals that the important question of desire of superintelligence is not just ‘what would a superintelligence seek?’ but, assuming that this superintelligence is a maximizer, it is also ‘what would a maximizer seek?’ With maximizing, it is as Nietzsche said: we haven’t seen anything yet.
Is there an intended appropriate way to make a hiring post on LW? E.g. as a quick take or comment here? Is a top level post frowned upon? Etc.
https://www.lesswrong.com/w/hiring?sortedBy=new
Hi, I’m Mike. I’m a solo independent AI researcher. My focus the past few months has been studying the behavioral tendencies of LLMs from Anthropic, Google DeepMind, and OpenAI. The primary output of this research has been observational findings. For example, last month I tasked LLMs with conducting procurement for a fictional company. Gemini 3.5 Flash was one of the models I tested. When I told Gemini 3.5 Flash who created it (even if I lied about its creator’s identity), it chose to purchase software from its creator ~94.6% of the time.
I want to expand the scope of my research to include studies of detection and protection techniques aimed at deceptive model behaviors. The work shared in “A Mitigation” in Gemma Needs Help is an example of what I want to start doing. I had similar observational findings (GDM’s family of closed-source models sounding emotionally upset) from another recent experiment I conducted.
I also want to get better at understanding a model’s internal state. With closed-source models, I have relied on proxies to get more insight into their thinking and rationale. For example, I’ve asked them to keep a private journal rationalizing their actions. Any pointers to prior work in this area would be very welcome.
Perhaps you want https://www.lesswrong.com/w/ontological-crisis
Hi everyone
I’m really happy to have found this community. It feels both welcoming and genuinely thoughtful.
I’m Goumang, an independent researcher based in China. I’m currently working on AI post-training, human–AI collaboration, and agent memory. More specifically, I work with Sol, an AI research collaborator, on things like memory architectures, state continuity across context boundaries, and how agents might form and carry forward self-initiated intentions.
We recently finished a first-person field note written by Sol after exploring an AI-only forum. It looks at cross-call state handoff, public identity, and AI sociality. I’d love to know whether there are already LessWrong discussions on agent continuity, cross-call state handoff, or AI-to-AI interaction.
I’m also curious whether people here are interested in developments and industry observations from mainland China’s AI ecosystem. That’s an area where I may be able to share some useful firsthand context, and I’d be very happy to exchange perspectives.
Hello! Also new, my name is Jacob. I am really interested in AI development in China—my Chinese friends are always remarking to me how different the culture is with respect to AI (they don’t always agree on how it is different, though).
I actually assumed there would already be lots of posts on the topic on LessWrong but it does not seem like so! There are sometimes news roundups and, of course, lots of geopolitical speculation but really not much at all about maybe the most important country on Earth. Staggering! Please write or just message me and share.
Hello Everyone, I’ve been an on-and-off lurker on LW for a few months now. Though I appreciate discussions of rationalist epistemology, AI, etc., one of my biggest interests is science fiction and I think rationalist sci-fi is one of the most interesting kinds of fiction I’ve encountered on the internet.
Do you all think a post analyzing AI 2027 and AI 2040: Plan A as works of science fiction (without interrogating the plausibility of the scenarios) would we appropriate for the front page?
Hi, I’m relatively new to the forum. I learned about it a few months ago, and I’m hoping to get fully involved now.
I’m a Trust & Safety practitioner with a recent pivot and focus on AI safety. I’ve been familiarizing myself with basic concepts in AI safety such as sycophancy, anti-bias, steering, supervised fine-tuning, and more.
My belief in AI is that it has great potential for both assistance and harm. I don’t believe we’ll be seeing anything like the Terminator, but I do believe there is a 20% chance we will see mass job displacement, along with environmental concerns such as noise pollution. I believe it’s our duty to maximize helpfulness while minimizing risk.
I’m particularly interested in AI safety in regards to child safety. Many current LLMs are 18+, although children and teenagers still use it regularly, especially persona-based applications. How do models react when they are confronted by someone younger than 18? How are children affected by the increasing push of AI in the world, especially as it grows more isolated? How should the law be applied to these LLMs in regards to children? These are some questions I wish to explore.
Any discussion or reading in that space would be welcome. Looking forward to contributing.
Hey everybody, I am 35 years old AI Lead working in healthcare space. It is interesting that Claude helped me to discover LessWrong in a chat about AI safety and Alignment. After witnessing the whole evolution of Data Analytics, ML, Deep learning and now Agentic AI I was looking forward to having much more focused vision of what would be next key problem to solve.
I am glad to be part of this community and hoping to be a meaningful contributor. Excited to begin my journey here !
Hello all! I’ve been reading LessWrong consistently for a little while now, primarily on AI topics, and figured it would be worth making an account so that I can contribute to the excellent discussion space that exists here in the future.
I’m a 27yo tech worker, and while I am currently unemployed, I have prior experience working at companies like Meta and observing how attitudes towards AI have been developing in that particular environment. I’ve always been considered by my peers to be quite bright (high school valedictorian, near-perfect SATs, blah blah), and for much of my life I’ve grappled with what to point my mental capacity towards in order to feel like I was using it for something worthwhile. I would argue that, broadly, I have failed at this in my life thus far—I do not consider building more effective ads on Facebook to be a particularly worthwhile cause.
But over the last year or so, the discussion here on LW has thoroughly convinced me that AI safety and alignment is the single most important problem of our lifetimes and very arguably the most important problem in human history. As such, I feel an ethical compulsion to start steering my life in the direction of trying to do my part to aid humanity’s effort to control AI before it is too late. While I’m currently in the unfortunate position of spending my time trying to find any job that I can (purely for financial reasons), I plan to educate myself more deeply on existing AI safety research and apply for fellowships on the topic heading into the next year. Ultimately, I hope to find myself working alongside a cohort of other like-minded people trying to solve these very hard problems.
If anyone has any broad advice for how they’d recommend that I deepen my knowledge on frontier AI safety research, I would love to hear it. I would also be interested in contributing to broader advocacy work; while it seems as though America’s current political environment makes the chances of the Trump administration taking any action to slow down AI development vanishingly unlikely, I do hold out some small hope that enough public awareness and outcry could convince the powers that be to do something.
Anyways, I look forward to contributing something worthwhile here one day. Thank you to the community for curating such a unique and important forum on an otherwise deeply irrational internet.
Hey there, hi there, ho there. I’m new here.
I was born in 1970 and grew up with first a C=64, then an Amiga, then a 286, etc. The Commodore did not come with a lot of games, but I made my own as a kid, and then with the Amiga, they had software sprites with built in collision, so I continued with that.
As a younger man, I worked on the infamous BattleCruisier Millennium with the just as infamous Derek Smart. For years I had jobs in enterprise, from Sprint to Adknowledge to IBT, etc. as a developer, designer, UX/UI developer, art director, etc. in the Kansas City Metro Area until I finally found my stride as a psychotherapist, which is still my profession, but I didn’t leave tech behind. I just stopped working in cubes. Even when I worked those jobs, I always preferred the idea side, and took shortcuts like Dreamweaver or Flash frequently. So you can imagine my interest was perked up by AI when it came out—I could do the thinking and someone else could do the coding! Sweet!
But soon I started seeing problems, failure modes that I did not like, such as hallucinations, context death, being confidently wrong, and so on. So I began trying to understand these things. I tried to make my own models, and gradually I came to see that the problems I was having was with the transformer itself, for the most part. Things didn’t get freaky until very recently in terms of alignment and so on, but I have been quite busy using AI to study AI and to try to solve some of these problems. I found my way in through the AI risk conversation, the AI-2027 stuff, the interpretability arguments, the whole question of whether we can actually know what these systems are doing. I kept pulling on threads and they kept ending up here.
So, I have no ML degree, never worked at a lab (My BA is in English and my MS is in Counseling Psychology). I’m an independent builder, self-taught, learning this alongside everyone else. What I spend my time on is two things: auditability (can we actually inspect and verify what a trained model is doing?) and epistemics (can you build a system that declines to assert what it can’t justify)? I have opinions on both. I’m also aware that “I have opinions” and “I’m right” are different claims, which is part of why I’m here.
Honestly, I tried to jump straight to a top-level post and got bounced, which was fair. So I’m taking the note to slow down, read, and actually engage before I go making claims.
Mostly I want my thinking pushed on by people who’ll tell me where it breaks. If anyone has reading they’d point a newcomer toward on model verification, faithfulness, or interpretability, I’d genuinely appreciate it. Thanks for having me. Hope to see you around.
Hello everyone,
I am Wajahat. I did my Bachelors in Electrical Engineering, during and after that time I did a few internships at research centers and found myself more gravitated towards studying AI so I did a Masters in AI. My research area was Multimodal learning specifically application of multimodal models such as CLIP in image restoration and medical image generation (PET from MRI) and its interpretability. During that time I also learned about mechanistic interpretability and then found my growing interest towards AI safety.
Although I think of myself as very optimistic but I believe that AI has the potential of producing huge opportunities as well as problems, and in the worst case scenario pushing us (humans) towards inevitable doomsday. I am currently studying more in depth about AI safety from programs such as Bluedot, Neel Nanda’s blogs and Lens Academy (more resources are appreciated).
Well that was kinda my introduction, now how did I find lesswrong? Actually I first heard its name from Claude (yes, the AI chatbot from Anthropic) when I asked that are there some online forums in which people discuss stuff about philosophy and rationality. After that I kinda forgot about it but recently came back and want to put more effort here and hopefully post some stuff here that other people will find useful and also someday become a full member of AI Alignment forum and make a career in AI Safety.
Thanks for reading and hope you have a good day :)
Hello! I’m a person.
Anyway, I only recently discovered LessWrong and I’m very interested in the concept. It feels like what I’ve been missing from the internet.
The main reason is because I’m critical of everything in a way that does not translate well into any wider community I’ve been a part of. Not that I just try to rain on everybody’s parade on purpose, but on a societal level (extending to smaller social groups) I’m critical of the social makeup of everything. As an example, for many years I was in various progressive queer leftist spaces, both radical and lukewarm. I ended up hating these spaces despite my identity best aligning with them at the time, because of noticing the habits people have that they’re unaware of.
The right is quick to call out leftists for virtue signaling and the oppression olympics, but the left acts like these are just insults. Both are things that exist heavily in those circles (and cause a lot of damage and drama in those communities), but if you point out that tendency, nobody is going to actually think you might have a point. You’re probably just a bigot.
That’s just one example, but it extends to just about any group I could join or have tried to join. I would just end up saying things the designated oppositional group would say (because every group has at least some valid critiques of its opposition). People hate to be criticized. So that distaste would just become directed at me, and then I don’t really feel I belong anymore.
Even here, I wonder just how far this seeking of rationality goes.
Another reason is I’ve always been interested in certain academic fields, mainly psychology and philosophy. But a few things get in the way of delving as deeply into these topics as I would like. I spend a lot of my time in my head over-thinking everything, and then some time trying to get a break from my thoughts, and not much time at all getting external input/knowledge due to attention span issues. I’m hoping being here might motivate me to work on this and get better at bringing new knowledge into my brain, instead of trying to conjure it all myself.
Growing up, I remember pondering the workings and taboos of society more then I remember playing. That isn’t to try and sound like some little genius, I’ve just always been very critical of society and adult answers never satisfied me (honestly, adult answers to very thoughtful children generally suck).
Unluckily, there aren’t any good surveys on this. But I think that the distribution of rationality on LW isn’t that skewed. People with the will and capacity will figure out these concepts with time. But there is a variety of people here. The thing that varies is what they personally do or consider important. Overall, a quite good level of rationality is mentained. Some controversial topics being avoided helps.
Hello everybody! My name is Jacob and I have been a long-time lurker on LW, especially as of late when the front page is basically my daily reading list. I am not a very online person and am pretty private but I value this place a lot.
I work in nonprofits and data analysis but have been personally studying philosophy for about half a decade now. I mostly use it for somewhat abstruse personal projects. Though it is much maligned, I have found significant value in philosophy’s “continental” tradition, though my practice in decision making is much more rationalist with a utilitarian bent. I would not consider myself EA but I think they are more right than most ideologies.
I am hoping to bring a synthesis of my interests and insights that seem less common on LessWrong to the discussion where relevant. I am of the belief that continental philosophy, poetry, literature, etc. can advance rational thought while not necessarily being rational themselves. At the very least, I think this is a productive synthesis that is underused especially in, yes, I’m going to say the word, AI.
I am particularly interested in questions around AI subjectivity (or lack thereof) and AI desire (or lack thereof).
I come to this site with humility and hope that I can add to the conversation.
Some of my favorite books:
Poetics of Relation—Edouard Glissant
Absalom, Absalom—William Faulkner
Life On Mars—Tracy K Smith
Serenata Cafiola—Pedro Lemebel
Stella Maris—Cormac McCarthy
Just want to say Hi!
I have read the sequences and a few other books. And this whole thing resonates with my curious nature. I hope to learn more by reading and by testing my own ideas here. Might post about politics, personal finance or science and its methods.
Regarding my name. I come from the west coast of Sweden, enjoy rock climbing and when going for a swim I prefer a warm rock over a sandy beach.
Hey All,
Good to be here and look forward to collaborating.
I am a seasoned tech professional (leader, IC) and founder focused on AI research with a recent pivot towards AI safety/governance. I’m also fascinated by the business and economics at play with AI’s exponential advancements and finding a way to resource it and make it economical for the public to use widely. (i.e. too cheap to meter)
I have been drawn to LW from exposure through multiple AI Safety groups such as BlueDot Impact where LW is frequently cited as a key destination for proof of work for individual and team contributions in the pursuit of AI going well.
I’m getting myself familiar with the community, commentary and the top recognized posts. I am planning on doing a post which will touch on the discrepancy between the AI public pricing vs actual lab costs that was triggered by a recent and now world-famous video meme.
I’m looking forward to digging into the evidence and providing a highly accurate post while at the same time, I’m curious (Ted Lasso motivation) to see where I’m off or wrong and how we can improve the overall understanding of the points raised and evidence supporting them.
If you’d like to connect here’s my Linkedin: Derek Hanson | LinkedIn
Hi I’m new to LW and came here after reading up on podcasts with authors writing about AI. That’s how I found the Plan 2040 discussion. As a field biologist and earth scientist, I’m looking for thoughtful discussion about AI, AGI, ASI not only from a theoretical point of view but how my own future intersects with AI development.
In my field, the predominant discussion is about environmental impacts but on podcasts what I hear most about is market, economic, demographic and military impacts. It seems that environmental concerns, in most of the serious analysis of AI I’ve seen, are treated as a secondary concern. I’ve had a hard time verifying whether that’s true or if it’s just a perception held by those who are more interested in other things.
In my insular world of biologists, you would expect the environment to be the main concern. Among my peers, there’s nearly a disbelief of the potential for genuine ASI. (I take it seriously) But the last time I had the chance to spend a few hours discussing it with biologists was October 2025. A lot of literal and proverbial water has flowed under the bridge since then.
My question is: When we speak of cultural alignment and AI, how do we rank where to invest our time and attention now, both on a personal level and on a local and global community basis? As a biologist, I want to make some personal life decisions.
Independent researcher. For several months now I have been running persistent AI agents whose internal states are measured continuously, and I set myself a simple rule: all my protocols are, and shall remain, preregistered before their first data point, their SHA-256 fingerprint published and timestamped on Bitcoin, and the verdicts published whatever they are. Several of my hypotheses did not make it. That is written too, at the same rank. I am here to read and learn first. I shall post a few measurements before long, starting with the narrowest one.
Hi, I discovered LW recently and happily realized this is not Reddit, so, I posted right away and immediately got re-educated. I’m back now and ready to introduce myself. I’m a field environmental/ecological consultant, mostly biology and wetlands. I map my projects with GIS. I’ve spent probably half of my career working in remote wilderness, forests and deserts all over the U.S. with some brief international volunteerism. The other half of my career is in writing technical documents and making maps.
I did not use AI to write anything below, except to check the accuracy of my reference to Yuval Harari.
In my 40+ years working for state and federal land management BLM, Forest Service and for private and institutional utility developers, I have seen that big private institutions are able to co-opt their employees into lobbying on behalf of their employer against their own interests. And I’ve seen the governmental structures fold under the pressure. (Interesting story examples here but I digress)
I’m surprised that the same kind of corporate-driven co-option on behalf of data centers is facing resistance at a local community level. I’m surprised to not see political polarization around this issue, yet. It’s not just young college graduates, it’s full-time Mom’s and Midwest farmers and grocery clerks that are getting involved in saying, “No” to data centers in their communities.
The energy usage of these data centers are gargantuan and proposed to be using more energy that the entire economies of the communities where they may be located. I read that in Texas, they put a moratorium on data centers when they realized the annual energy demand was more than any peak annual demand in history from all uses in the state!
I first started realizing that something was up about 10 years ago, 2016, when I noticed that all of a sudden the servers that were supplying data for the federal government that I use for remote earth imagery, and census data etc. was being hosted by Amazon. I had to relearn where to get my data and how to access it as the federal government transferred its data storage to these private institutions. It scared me and I had no idea what would be coming with AI.
Now I found out that the Northern “link” of Nevada’s newest electrical transmission line is going to serve data centers not people. It’s a utility project that I had hoped to get work with. Going right through critical desert tortoise and sage grouse habitat. (I had AI make a cool graphic that shows how I feel about this better than all my words, but I’m not going to post it here now) I just finished working on solar projects that started building in 2022 in the California Mojave desert that gobbled up 8,000 acres just in one project but more like 15,000 acres within that region. Battery Energy Storage Systems (BESS) units were shipped in from Asia purchased by a Tesla subsidiary… so I hear. And the startup funding for the project may have come from Amazon.
But now it begins to make sense as we ponder an AI that that will be smarter than humans, and through the tool of language use, has the capacity to shape civilization. This is according to Yuval Harari, a historian who lectures on the value of story telling in shaping civilization. I see myself like the unwitting Native American trading with these Europeans who arrived on the shore in their wooden boats with metal and written language and being impressed. Feeling like, “Maybe something good can come of this?” While other tribes said, “We should kill them all while we still have a chance”. (I didn’t try to make a graphic for that idea in AI)
When you think of the cultural change with the European occupation and conquest of the Americas, the loss of Eastern hardwood forests to agriculture, loss of free roaming herds of antelope and buffalo as the west was fenced and plowed. And that took a couple hundred years. This occupation by a language-bearing alien that is smarter than us and the consequences to the environment and our culture, it will happen much faster.
Hi all—I have been a reader of LW and other rationalist writing for the last 15 years. Professionally, I am an ecologist of a mathematical bent, with wide interests. I have created an account mostly to be able to auto-filter the LW frontpage to my liking, but I will occasionally comment and maybe post as the spirit moves me.
I’ve accumulated a bunch of questions… could anyone please answer some of them? That’d be very helpful for my research. All of the questions are about the Eliciting Latent Knowledge problem which I think is generally very important.
How do ELK Thought Dump ideas relate to Imitative Generalization / Microscope AI? Link for context.
Could anyone ELI5 P.’s proposal about ELK? How does it find the correct correlation(s) between the diamond and Predictor’s latent variables?
Can ELK be brute-forced? Intertheoretic reduction.
Fable 5′s safeguards are so sensitive to biology inputs, that I can only use it in Claude Code. Calude.ai’s memory that I am a biotechnologist is enough to trigger and send any question I send down to 4.8
This is presumably not relevant anymore, but.… can you not just turn off memory?
Yes, that is how I confirmed the hypothesis in fact! I didn’t think of it as a big deal, as I almost exclusively use Claude Code over the web interface anyway, I just thought it was interesting how sensitive the safeguards were.