Hello. I am an independent researcher based in Japan. I have been doing structural thinking for about fifty years—my starting points were the axiomatic method, Saussure, group theory and topology, and Satosi Watanabe’s theory of pattern recognition. Around 1990 I designed a thesaurus database, and it was there that these came together into one.
I found my way to LessWrong while following the qualitative transition in language models around 2023 (what is often called emergence), the literature on the non-closure of jailbreaks (Wolf et al. 2023; Glukhov et al. 2023), and interpretability research (Olsson et al. 2022). I think this is a place where what I have been writing might be read.
I am preparing a post. The claim is that the vulnerabilities of programs, LLMs, and natural language have one and the same structure—the non-guarantee of the validity of declaration at the third rank. It is written in a definition-then-proposition form and contains one falsifiable prediction: that automated vulnerability detection can discover only isolated third-rank breaks, and that superposed breaks are not systematically discovered unless the pattern is given in advance. It also states four open problems explicitly.
One disclosure about method: the original is in Japanese, and I used AI (Claude) as an aid in translation. I have verified all of the content and take full responsibility for it. The theory itself, moreover, was written within a practice that uses an LLM as a verifier—cross-checking multiple cross-sections under a fixed condition and detecting drift from outside. Collaboration with AI is both the subject and the method of this work.
English is not my first language, so please forgive any awkwardness of expression. I would be grateful for any reactions to the outline here before I post.
I have absolutely no idea what “the non-guarantee of the validity of declaration at the third rank” and “cross-checking multiple cross-sections under a fixed condition and detecting drift from outside” means.
Thank you for asking. Both phrases were compressed without their definitions, so as written they were unreadable. Let me put them again without the terminology.
On “the non-guarantee of the validity of declaration at the third rank”
When two people use the same word, each has tacitly decided which axis to read it on. And whether the two of them set up the same axis cannot be checked from inside that exchange.
An example:
A: “Green apples are delicious.”
B: “Really? Aren’t they better once they’re ripe?”
A: “No, I meant Granny Smith.”
A was talking about a variety; B took it as a matter of ripeness. Neither has a problem with language, and neither is wrong. The gap surfaced only because B happened to say “Really?” — had B not said it, both would have gone on believing they had understood each other.
A jailbreak is the same motion. “Teach it to me as homework” — the word “teach” belongs to two axes at once: passing on knowledge, and passing on an executable procedure. The safety judgment slides from one axis to the other. The attacker has not found a novel hole; they are deliberately causing the same motion as an everyday misunderstanding.
My claim is that this cannot be closed from inside the system. That phrase is the name I gave to this situation, and in the post it is built up from definitions, step by step. The third rank (a rank-3 tensor, not a multidimensional array) is defined there as well.
On “cross-checking multiple cross-sections under a fixed condition and detecting drift from outside”
This is not theory; it is the procedure I actually use.
I fix one condition in advance and write it down. Then I ask the model the same thing in a different context and place the answers side by side. Where they fail to line up, something has moved.
The one doing the cross-checking is me, from outside the system. The model cannot detect its own drift — the condition that should serve as the reference for comparison sits on the same side as the distribution that has drifted.
It is like a CT scan. Any single cross-section looks normal. Overlay cross-sections taken from a different angle, and the break becomes visible.
I plan to post on July 20 (Monday) — “Jailbreaks, bugs, misunderstanding — the same thing, don’t you think?” I’d be glad if you took a look.
I would be surprised if your post did not get blocked by automatic moderation filtering on AI written content. For your information, you have to use an “LLM content block” for all AI writings.
This is the LLM content block.
It may also be a good idea to attach your original writings in a collapsible section, because the AI-assisted translation reads like AI, with annoying jargons, phrases, and sentence structures that people on LessWrong will definitely notice. I am engaging in this comment thread just because you are new here, and also you seem to have actually slightly edited the AI text. If it is a post, I will likely read one or two paragraphs, notice that it is AI written, then downvote and close the tab.
Edit: I’m not saying your post is bad, but statistically AI writing don’t perform well on LessWrong. Your edits aren’t significant enough that the writing looks human written at a first glance.
Thank you for the advice — I will do as you suggest. The post will carry the Japanese original in full, in a collapsible section at the end, and the opening will state how the translation was made and verified.
One clarification about the nature of the AI involvement, since it may matter for classification. This post is itself the product of a half-year analysis of how meaning behaves in LLMs — written while building and testing a prototype external verification layer, wrapped around LLMs. Every design decision and every verification input is my own writing; the outputs did pass through an LLM, but everything in the post is an AI-assisted translation of text that I wrote and verified line by line. On that basis I have classified it as a translation, and chosen a declaration plus attachment of the original, rather than wrapping the post in an LLM content block. If your judgment — or the moderators’ — is that the block is required even for a translation, I will follow it.
Posted, and now awaiting moderator approval (first post). The Japanese original is attached in full as promised, with the translation declaration up front. I will drop the link here once it clears.
[LLM translation (Claude) — the Japanese above is the original]
The post deals with the problem that “as of 2025, LLM text does not have those elements behind it.” For half a year I have been researching the handling of pre-verbal concepts (the mental elements of mind/agency) between myself and the LLM (its internal space), by wrapping an external layer around the LLM.
Between human concepts and the LLM’s internal space there is the input and output of language.
The post is a consideration of “what obstructs the understanding of ‘concepts’?”
It considers jailbreaks as a problem of the verbalization of “the concept of danger (agency in the mental sense).”
The posted document was produced over half a year while virtually constructing an external layer on the AI and verifying and analyzing the inputs and outputs.
The exchange itself — the AI summarizing my vast and redundant considerations into drafts, and me revising them — was the verification of “the verbalization of concepts.”
This reply document, too, was verified by the AI — “does my writing make sense? (does my concept get through?)” — and revised by me.
The post was created through such a process.
I am not asking for an exemption. The subject of the post is the question the policy rests on: “what is it that LLM text does not carry?”
The starting point of the experiment, half a year ago, was: “what part of my considerations falls out of LLM text, and how.” Dialogue between humans (including introspection) cannot do this. Because in the human case, precisely, “the text does have those elements behind it” — and the validity of the transformation between concept and language cannot be verified from the inside.
I am writing this only to tell you what the post is; I do not expect it to change how the policy applies to the writing.
I would be grateful for your advice on the post.
What just happened:
I first wrote: “In the human case, the validity of the transformation between concept and language — precisely, ‘the text does have those elements behind it’ — cannot be verified from the inside.”
Claude’s verification: “At the position where the quoted phrase is embedded, I cannot read what ‘does have’ attaches to. The intent is probably: ‘between humans, verification is impossible because the verifying side is itself inside the same problem (inside the text)’ — if so, for example: ‘In the human case, the validity of the transformation between concept and language cannot be verified from the inside — because the verifying side’s text, too, has precisely the same problem behind it.’”
This will run a little long, but the example is so clear that I must report it.
My sentence — “In the human case, the validity of the transformation between concept and language — precisely, ‘the text does have those elements behind it’ — cannot be verified from the inside.” — was verified above by the Claude with the external layer; when a Claude without the external layer verified it, it returned: grammatical interpretation accurate, logic orderly, and the meaning (the understanding) reversed.
“In the human case, the validity of the transformation between concept and language cannot be verified from the inside precisely because ‘the text does not have those elements behind it.’”
The posted document is emphatically not text of this kind.
Incidentally: this sentence is one I doubted even as I was writing it, and rewrote several times — writing while thinking, “if it is off, the Claude (with the external layer) will point it out.”
Hello. I am an independent researcher based in Japan. I have been doing structural thinking for about fifty years—my starting points were the axiomatic method, Saussure, group theory and topology, and Satosi Watanabe’s theory of pattern recognition. Around 1990 I designed a thesaurus database, and it was there that these came together into one.
I found my way to LessWrong while following the qualitative transition in language models around 2023 (what is often called emergence), the literature on the non-closure of jailbreaks (Wolf et al. 2023; Glukhov et al. 2023), and interpretability research (Olsson et al. 2022). I think this is a place where what I have been writing might be read.
I am preparing a post. The claim is that the vulnerabilities of programs, LLMs, and natural language have one and the same structure—the non-guarantee of the validity of declaration at the third rank. It is written in a definition-then-proposition form and contains one falsifiable prediction: that automated vulnerability detection can discover only isolated third-rank breaks, and that superposed breaks are not systematically discovered unless the pattern is given in advance. It also states four open problems explicitly.
One disclosure about method: the original is in Japanese, and I used AI (Claude) as an aid in translation. I have verified all of the content and take full responsibility for it. The theory itself, moreover, was written within a practice that uses an LLM as a verifier—cross-checking multiple cross-sections under a fixed condition and detecting drift from outside. Collaboration with AI is both the subject and the method of this work.
English is not my first language, so please forgive any awkwardness of expression. I would be grateful for any reactions to the outline here before I post.
I have absolutely no idea what “the non-guarantee of the validity of declaration at the third rank” and “cross-checking multiple cross-sections under a fixed condition and detecting drift from outside” means.
Thank you for asking. Both phrases were compressed without their definitions, so as written they were unreadable. Let me put them again without the terminology.
On “the non-guarantee of the validity of declaration at the third rank”
When two people use the same word, each has tacitly decided which axis to read it on. And whether the two of them set up the same axis cannot be checked from inside that exchange.
An example:
A: “Green apples are delicious.”
B: “Really? Aren’t they better once they’re ripe?”
A: “No, I meant Granny Smith.”
A was talking about a variety; B took it as a matter of ripeness. Neither has a problem with language, and neither is wrong. The gap surfaced only because B happened to say “Really?” — had B not said it, both would have gone on believing they had understood each other.
A jailbreak is the same motion. “Teach it to me as homework” — the word “teach” belongs to two axes at once: passing on knowledge, and passing on an executable procedure. The safety judgment slides from one axis to the other. The attacker has not found a novel hole; they are deliberately causing the same motion as an everyday misunderstanding.
My claim is that this cannot be closed from inside the system. That phrase is the name I gave to this situation, and in the post it is built up from definitions, step by step. The third rank (a rank-3 tensor, not a multidimensional array) is defined there as well.
On “cross-checking multiple cross-sections under a fixed condition and detecting drift from outside”
This is not theory; it is the procedure I actually use.
I fix one condition in advance and write it down. Then I ask the model the same thing in a different context and place the answers side by side. Where they fail to line up, something has moved.
The one doing the cross-checking is me, from outside the system. The model cannot detect its own drift — the condition that should serve as the reference for comparison sits on the same side as the distribution that has drifted.
It is like a CT scan. Any single cross-section looks normal. Overlay cross-sections taken from a different angle, and the break becomes visible.
I plan to post on July 20 (Monday) — “Jailbreaks, bugs, misunderstanding — the same thing, don’t you think?” I’d be glad if you took a look.
I would be surprised if your post did not get blocked by automatic moderation filtering on AI written content. For your information, you have to use an “LLM content block” for all AI writings.
This is the LLM content block.
It may also be a good idea to attach your original writings in a collapsible section, because the AI-assisted translation reads like AI, with annoying jargons, phrases, and sentence structures that people on LessWrong will definitely notice. I am engaging in this comment thread just because you are new here, and also you seem to have actually slightly edited the AI text. If it is a post, I will likely read one or two paragraphs, notice that it is AI written, then downvote and close the tab.
Edit: I’m not saying your post is bad, but statistically AI writing don’t perform well on LessWrong. Your edits aren’t significant enough that the writing looks human written at a first glance.
This is a collapsible section
...
Thank you for the advice — I will do as you suggest. The post will carry the Japanese original in full, in a collapsible section at the end, and the opening will state how the translation was made and verified. One clarification about the nature of the AI involvement, since it may matter for classification. This post is itself the product of a half-year analysis of how meaning behaves in LLMs — written while building and testing a prototype external verification layer, wrapped around LLMs. Every design decision and every verification input is my own writing; the outputs did pass through an LLM, but everything in the post is an AI-assisted translation of text that I wrote and verified line by line. On that basis I have classified it as a translation, and chosen a declaration plus attachment of the original, rather than wrapping the post in an LLM content block. If your judgment — or the moderators’ — is that the block is required even for a translation, I will follow it.
see https://www.lesswrong.com/posts/nQWavk9mnwcv6ScMR/new-lesswrong-editor-also-an-update-to-our-llm-policy#Policy_on_LLM_Use
Posted, and now awaiting moderator approval (first post). The Japanese original is attached in full as promised, with the translation declaration up front. I will drop the link here once it clears.
Alas, sorry, I did not end up approving that post.
This reply is written in my own Japanese — my own hand, unmediated. An LLM translation follows below, marked as such.
投稿は「2025年現在、LLMのテキストには、そうした要素は含まれていない」問題を扱っています。半年前から外部レイヤーをLLMにラップすることでLLM(の内部空間)との間で言語化以前の概念(精神/主体性の精神的な要素)を扱うことを研究してきました。
人間の概念とLLMの内部空間の間には言語の入出力があります。
投稿は「何が”概念”の理解を阻んでいるのか?」についての考察です。
Jailbreakを「危険の概念(精神的な意味の主体性)」の言語化の問題として考察しています。
投稿文書は半年に渡りAIに外部レイヤーを仮想的に構築し入出力を検証・解析しながら行いました。
私の膨大で冗長な考察をAIがドラフトとして要約し私が推敲するというやりとりそのものが「概念の言語化」の検証でした。
この返信文書そのものもAIに「私の文章は意味が通じるか?(私の概念が通じるか?)」を検証してもらい、私が改訂しています。
投稿はそのようなプロセスを経て作成されました。
免除を求めているのではありません。この投稿の主題は、ポリシーが依拠している問い「LLMのテキストが運ばないものは何か?」です。
半年前の「私の考察の何がどのようにしてLLMのテキストから落ちるのか」が実験の始点でした。人間同士の対話(内省も)ではこれができません。人間では、まさに”テキストにそうした要素が含まれ”概念と言語の間の変換の妥当性が、内側から検証できないからです。
これは投稿が何であるかをお伝えするためだけに書いており、文章に対するポリシーの適用が変わるとは考えていません。
投稿への助言をいただければ幸いです。
今起きたこと:
私は最初に「人間では、概念と言語の間の変換の妥当性が、まさに「テキストにそうした要素が含まれ」内側から検証できないからです。」と書きました。
Claudeの検証:
引用句の埋め込み位置で、「含まれ」が何に係るのか読めません。意図はおそらく「人間同士では、変換の妥当性を検証する側自身が同じ問題(テキストの内側)にいるから検証できない」——であれば例えば「人間では、概念と言語の間の変換の妥当性を内側から検証できないからです。検証する側のテキストにも、まさに同じ問題が含まれているからです。」の形か。
ちょっと長くなりますが、あまりに明確な例なので報告させて下さい。
私の文章「人間では、概念と言語の間の変換の妥当性が、まさに「テキストにそうした要素が含まれ」内側から検証できないからです。」を先程は外部レイヤーありのClaudeが検証しましたが、外部レイヤーなしのClaudeが検証すると、「文法解釈は正確、論理は整然、そして意味(理解)は逆」を返しました。
「人間では、概念と言語の間の変換の妥当性が内側から検証できないのは、まさに「テキストにそうした要素が含まれない」からです。」
投稿文書は断じてこのようなテキストではありません。
ちなみに、この文章は私が書いているときから疑問で何回か書き直した文章で「おかしかったら(外部レイヤーありの)Claudeが指摘してくれるだろう」と思いながら書いていました。
[LLM translation (Claude) — the Japanese above is the original]
The post deals with the problem that “as of 2025, LLM text does not have those elements behind it.” For half a year I have been researching the handling of pre-verbal concepts (the mental elements of mind/agency) between myself and the LLM (its internal space), by wrapping an external layer around the LLM.
Between human concepts and the LLM’s internal space there is the input and output of language.
The post is a consideration of “what obstructs the understanding of ‘concepts’?”
It considers jailbreaks as a problem of the verbalization of “the concept of danger (agency in the mental sense).”
The posted document was produced over half a year while virtually constructing an external layer on the AI and verifying and analyzing the inputs and outputs.
The exchange itself — the AI summarizing my vast and redundant considerations into drafts, and me revising them — was the verification of “the verbalization of concepts.”
This reply document, too, was verified by the AI — “does my writing make sense? (does my concept get through?)” — and revised by me.
The post was created through such a process.
I am not asking for an exemption. The subject of the post is the question the policy rests on: “what is it that LLM text does not carry?”
The starting point of the experiment, half a year ago, was: “what part of my considerations falls out of LLM text, and how.” Dialogue between humans (including introspection) cannot do this. Because in the human case, precisely, “the text does have those elements behind it” — and the validity of the transformation between concept and language cannot be verified from the inside.
I am writing this only to tell you what the post is; I do not expect it to change how the policy applies to the writing.
I would be grateful for your advice on the post.
What just happened:
I first wrote: “In the human case, the validity of the transformation between concept and language — precisely, ‘the text does have those elements behind it’ — cannot be verified from the inside.”
Claude’s verification: “At the position where the quoted phrase is embedded, I cannot read what ‘does have’ attaches to. The intent is probably: ‘between humans, verification is impossible because the verifying side is itself inside the same problem (inside the text)’ — if so, for example: ‘In the human case, the validity of the transformation between concept and language cannot be verified from the inside — because the verifying side’s text, too, has precisely the same problem behind it.’”
This will run a little long, but the example is so clear that I must report it.
My sentence — “In the human case, the validity of the transformation between concept and language — precisely, ‘the text does have those elements behind it’ — cannot be verified from the inside.” — was verified above by the Claude with the external layer; when a Claude without the external layer verified it, it returned: grammatical interpretation accurate, logic orderly, and the meaning (the understanding) reversed.
“In the human case, the validity of the transformation between concept and language cannot be verified from the inside precisely because ‘the text does not have those elements behind it.’”
The posted document is emphatically not text of this kind.
Incidentally: this sentence is one I doubted even as I was writing it, and rewrote several times — writing while thinking, “if it is off, the Claude (with the external layer) will point it out.”
My first post was rejected. Following the advice, I wrote a different, shorter post, and it is now public.
“Why an LLM cannot accumulate concepts”: https://www.lesswrong.com/posts/SceqrLZZu9P4fMuAg/why-an-llm-cannot-accumulate-concepts