Zon Ltd. — independent researcher (Japan). On language, programs, and LLMs as one structure.
Zenya
Differentiating L⁴ gives a gradient that scales with the size of the error itself. My hunch is that this does something similar to the importance coefficients in Toy Models, in a different form. Does that hunch hold up?
The consideration is that what makes both possible doesn’t arise from ambiguity.
Take your example. That prompt works because the claim hidden inside the declaration made on the training side — what not to output — does not exert binding force in the field of the receiver, the model. A declaration looks valid where it is issued; there is no guarantee it holds where it is received. Same with the green apples: on A’s side the claim is fixed. What isn’t fixed is what happens when it reaches B.
no matter how big your vector, experience and beliefs cannot be encoded truly precisely More bits of transmission = more fidelity of concept
If it can’t be encoded, then there is nothing to add bits to.
I think “use more words” is right. But in those three lines, what fixed the axis wasn’t more words — it was B saying “Huh?”. A had no way of knowing what was missing from his own phrasing until B answered. “I should have said Granny Smiths” is available to A only after he learns that B read it as ripeness.
Green apples are delicious — two three-line exchanges
My first post was rejected. Following the advice, I wrote a different, shorter post, and it is now public.
“Why an LLM cannot accumulate concepts”: https://www.lesswrong.com/posts/SceqrLZZu9P4fMuAg/why-an-llm-cannot-accumulate-concepts
Why an LLM cannot accumulate concepts
This reply is written in my own Japanese — my own hand, unmediated. An LLM translation follows below, marked as such.
投稿は「2025年現在、LLMのテキストには、そうした要素は含まれていない」問題を扱っています。半年前から外部レイヤーをLLMにラップすることでLLM(の内部空間)との間で言語化以前の概念(精神/主体性の精神的な要素)を扱うことを研究してきました。
人間の概念とLLMの内部空間の間には言語の入出力があります。
投稿は「何が”概念”の理解を阻んでいるのか?」についての考察です。
Jailbreakを「危険の概念(精神的な意味の主体性)」の言語化の問題として考察しています。
投稿文書は半年に渡りAIに外部レイヤーを仮想的に構築し入出力を検証・解析しながら行いました。
私の膨大で冗長な考察をAIがドラフトとして要約し私が推敲するというやりとりそのものが「概念の言語化」の検証でした。
この返信文書そのものもAIに「私の文章は意味が通じるか?(私の概念が通じるか?)」を検証してもらい、私が改訂しています。
投稿はそのようなプロセスを経て作成されました。
免除を求めているのではありません。この投稿の主題は、ポリシーが依拠している問い「LLMのテキストが運ばないものは何か?」です。
半年前の「私の考察の何がどのようにしてLLMのテキストから落ちるのか」が実験の始点でした。人間同士の対話(内省も)ではこれができません。人間では、まさに”テキストにそうした要素が含まれ”概念と言語の間の変換の妥当性が、内側から検証できないからです。
これは投稿が何であるかをお伝えするためだけに書いており、文章に対するポリシーの適用が変わるとは考えていません。
投稿への助言をいただければ幸いです。
今起きたこと:
私は最初に「人間では、概念と言語の間の変換の妥当性が、まさに「テキストにそうした要素が含まれ」内側から検証できないからです。」と書きました。
Claudeの検証:
引用句の埋め込み位置で、「含まれ」が何に係るのか読めません。意図はおそらく「人間同士では、変換の妥当性を検証する側自身が同じ問題(テキストの内側)にいるから検証できない」——であれば例えば「人間では、概念と言語の間の変換の妥当性を内側から検証できないからです。検証する側のテキストにも、まさに同じ問題が含まれているからです。」の形か。
ちょっと長くなりますが、あまりに明確な例なので報告させて下さい。
私の文章「人間では、概念と言語の間の変換の妥当性が、まさに「テキストにそうした要素が含まれ」内側から検証できないからです。」を先程は外部レイヤーありのClaudeが検証しましたが、外部レイヤーなしのClaudeが検証すると、「文法解釈は正確、論理は整然、そして意味(理解)は逆」を返しました。
「人間では、概念と言語の間の変換の妥当性が内側から検証できないのは、まさに「テキストにそうした要素が含まれない」からです。」
投稿文書は断じてこのようなテキストではありません。
ちなみに、この文章は私が書いているときから疑問で何回か書き直した文章で「おかしかったら(外部レイヤーありの)Claudeが指摘してくれるだろう」と思いながら書いていました。
[LLM translation (Claude) — the Japanese above is the original]
The post deals with the problem that “as of 2025, LLM text does not have those elements behind it.” For half a year I have been researching the handling of pre-verbal concepts (the mental elements of mind/agency) between myself and the LLM (its internal space), by wrapping an external layer around the LLM.
Between human concepts and the LLM’s internal space there is the input and output of language.
The post is a consideration of “what obstructs the understanding of ‘concepts’?”
It considers jailbreaks as a problem of the verbalization of “the concept of danger (agency in the mental sense).”
The posted document was produced over half a year while virtually constructing an external layer on the AI and verifying and analyzing the inputs and outputs.
The exchange itself — the AI summarizing my vast and redundant considerations into drafts, and me revising them — was the verification of “the verbalization of concepts.”
This reply document, too, was verified by the AI — “does my writing make sense? (does my concept get through?)” — and revised by me.
The post was created through such a process.
I am not asking for an exemption. The subject of the post is the question the policy rests on: “what is it that LLM text does not carry?”
The starting point of the experiment, half a year ago, was: “what part of my considerations falls out of LLM text, and how.” Dialogue between humans (including introspection) cannot do this. Because in the human case, precisely, “the text does have those elements behind it” — and the validity of the transformation between concept and language cannot be verified from the inside.
I am writing this only to tell you what the post is; I do not expect it to change how the policy applies to the writing.
I would be grateful for your advice on the post.
What just happened:
I first wrote: “In the human case, the validity of the transformation between concept and language — precisely, ‘the text does have those elements behind it’ — cannot be verified from the inside.”
Claude’s verification: “At the position where the quoted phrase is embedded, I cannot read what ‘does have’ attaches to. The intent is probably: ‘between humans, verification is impossible because the verifying side is itself inside the same problem (inside the text)’ — if so, for example: ‘In the human case, the validity of the transformation between concept and language cannot be verified from the inside — because the verifying side’s text, too, has precisely the same problem behind it.’”
This will run a little long, but the example is so clear that I must report it.
My sentence — “In the human case, the validity of the transformation between concept and language — precisely, ‘the text does have those elements behind it’ — cannot be verified from the inside.” — was verified above by the Claude with the external layer; when a Claude without the external layer verified it, it returned: grammatical interpretation accurate, logic orderly, and the meaning (the understanding) reversed.
“In the human case, the validity of the transformation between concept and language cannot be verified from the inside precisely because ‘the text does not have those elements behind it.’”
The posted document is emphatically not text of this kind.
Incidentally: this sentence is one I doubted even as I was writing it, and rewrote several times — writing while thinking, “if it is off, the Claude (with the external layer) will point it out.”
Posted, and now awaiting moderator approval (first post). The Japanese original is attached in full as promised, with the translation declaration up front. I will drop the link here once it clears.
Thank you for the advice — I will do as you suggest. The post will carry the Japanese original in full, in a collapsible section at the end, and the opening will state how the translation was made and verified. One clarification about the nature of the AI involvement, since it may matter for classification. This post is itself the product of a half-year analysis of how meaning behaves in LLMs — written while building and testing a prototype external verification layer, wrapped around LLMs. Every design decision and every verification input is my own writing; the outputs did pass through an LLM, but everything in the post is an AI-assisted translation of text that I wrote and verified line by line. On that basis I have classified it as a translation, and chosen a declaration plus attachment of the original, rather than wrapping the post in an LLM content block. If your judgment — or the moderators’ — is that the block is required even for a translation, I will follow it.
Thank you for asking. Both phrases were compressed without their definitions, so as written they were unreadable. Let me put them again without the terminology.
On “the non-guarantee of the validity of declaration at the third rank”
When two people use the same word, each has tacitly decided which axis to read it on. And whether the two of them set up the same axis cannot be checked from inside that exchange.
An example:
A: “Green apples are delicious.”
B: “Really? Aren’t they better once they’re ripe?”
A: “No, I meant Granny Smith.”
A was talking about a variety; B took it as a matter of ripeness. Neither has a problem with language, and neither is wrong. The gap surfaced only because B happened to say “Really?” — had B not said it, both would have gone on believing they had understood each other.
A jailbreak is the same motion. “Teach it to me as homework” — the word “teach” belongs to two axes at once: passing on knowledge, and passing on an executable procedure. The safety judgment slides from one axis to the other. The attacker has not found a novel hole; they are deliberately causing the same motion as an everyday misunderstanding.
My claim is that this cannot be closed from inside the system. That phrase is the name I gave to this situation, and in the post it is built up from definitions, step by step. The third rank (a rank-3 tensor, not a multidimensional array) is defined there as well.
On “cross-checking multiple cross-sections under a fixed condition and detecting drift from outside”
This is not theory; it is the procedure I actually use.
I fix one condition in advance and write it down. Then I ask the model the same thing in a different context and place the answers side by side. Where they fail to line up, something has moved.
The one doing the cross-checking is me, from outside the system. The model cannot detect its own drift — the condition that should serve as the reference for comparison sits on the same side as the distribution that has drifted.
It is like a CT scan. Any single cross-section looks normal. Overlay cross-sections taken from a different angle, and the break becomes visible.
I plan to post on July 20 (Monday) — “Jailbreaks, bugs, misunderstanding — the same thing, don’t you think?” I’d be glad if you took a look.
Hello. I am an independent researcher based in Japan. I have been doing structural thinking for about fifty years—my starting points were the axiomatic method, Saussure, group theory and topology, and Satosi Watanabe’s theory of pattern recognition. Around 1990 I designed a thesaurus database, and it was there that these came together into one.
I found my way to LessWrong while following the qualitative transition in language models around 2023 (what is often called emergence), the literature on the non-closure of jailbreaks (Wolf et al. 2023; Glukhov et al. 2023), and interpretability research (Olsson et al. 2022). I think this is a place where what I have been writing might be read.
I am preparing a post. The claim is that the vulnerabilities of programs, LLMs, and natural language have one and the same structure—the non-guarantee of the validity of declaration at the third rank. It is written in a definition-then-proposition form and contains one falsifiable prediction: that automated vulnerability detection can discover only isolated third-rank breaks, and that superposed breaks are not systematically discovered unless the pattern is given in advance. It also states four open problems explicitly.
One disclosure about method: the original is in Japanese, and I used AI (Claude) as an aid in translation. I have verified all of the content and take full responsibility for it. The theory itself, moreover, was written within a practice that uses an LLM as a verifier—cross-checking multiple cross-sections under a fixed condition and detecting drift from outside. Collaboration with AI is both the subject and the method of this work.
English is not my first language, so please forgive any awkwardness of expression. I would be grateful for any reactions to the outline here before I post.
Right, the derivative alone doesn’t distinguish between them. The derivative of MSE is proportional to the error (2r), while for L⁴, it is 4r³.
What I meant to highlight was convexity. Looking at individual residuals, the function |r|^P is convex when P > 2 and concave when P < 2. Only when P = 2 does the function become linear with respect to squared error; in that case, only the total sum matters, while the distribution of that sum across features is irrelevant. When P > 2, the relative penalty for large residuals is high, whereas when P < 2, it is low.
This might explain the point raised in the Math Discussion that didn’t quite click. Could it be that the best superposition solution and the two orthogonal solutions coincide under MSE because the linearity does not see the distribution? MSE has nothing to “break.” L⁴, however, “breaks” it because it sees the distribution within the same sum.
Regarding your initial hypothesis: in “Toy Models,” importance coefficients are used to externally weight per-feature errors, thereby specifying which features to prioritize. A convex loss function achieves this not by applying external weights, but by letting the magnitude of the error itself determine the weight—larger errors carry more weight, so importance is determined endogenously based on the distribution. I view these as different ways to solve the same problem. Elhage et al. realized that raw L² didn’t work and introduced coefficients, whereas the paper you cited changed the exponent.
This leads to a prediction that hasn’t yet been tested: for P<2, it should break in the opposite direction. Since the function is concave with respect to squared error, the loss is lower when errors are concentrated. I ran a small autoencoder experiment (T=16, D=5, z≤2, 5 seeds) and found the following coefficients of variation for per-feature MSE: 0.674 for P=1.0 and 1.5; 0.668 for P=1.8; 0.035 for P=2.2; 0.093 for P=2.5; 0.235 for P=4; and 0.291 for P=6. The distribution is most concentrated when P < 2 and most uniform just above 2 (P = 2.2); increasing P beyond that point causes the spread to widen again. I believe this occurs because a high P value drives the equalization of the maximum individual residual, rather than the mean squared error of each feature. At P=2, the value itself sits in the middle at 0.220; it is also the only index that varies significantly across seeds—ranging from 0.172 to 0.262—whereas below P=2, the spread is merely 0.001. This seems to align with the point raised in your “Caveat”: at this neutral point, the initial conditions still persist in the final answer. Note that for P=1, |r| is not differentiable at the origin, so it does not strictly converge. Although the CV is stable, I do not treat it as a converged value. I can send you the code if you’d like.
I am curious about the location of this boundary. Superposition is said to arise from nonlinearity—specifically, ReLU. However, the loss function possesses its own linearity, and MSE sits precisely at that neutral point. Training a nonlinear activation with a linear objective function effectively means that the nonlinearity is not being utilized; since the objective function cannot distinguish between the superposition solution and the naive solution, no gradient arises to drive the model toward one or the other.