I want to make better discoveries about a research-project topic to use for the comprehensive selection admissions process of the School of Computing at Institute of Science Tokyo (a Japanese university).
Background
Considering the distribution of Less Worng users, there do not seem to be very many Japanese users, so in order to explain the background, I need to give at least a minimal explanation of the Japanese university admissions system.
First, in order to apply to a Japanese national university, students first take something called the Common Test in January. Then, in February, they take an individual university examination (English, mathematics, chemistry, and physics). Some people may wonder why there are two examinations. The Common Test is administered by the national government, whereas the individual examinations are created by each university, and their difficulty and tendencies differ greatly between universities. Each university converts the Common Test score according to a certain weighting, and admission is decided based on the sum of that score and the score from the university’s individual examination.
However, each university also has distinctive admissions methods separate from this. For example, in the admissions method I am considering, the first-stage selection is based only on the Common Test and the evaluation of a research project conducted before the admissions process, and the second-stage selection is based only on an interview of about 30 minutes.
The advantage of this method is obvious. Compared with the former method, the required academic level is relatively lower. Under the ordinary method, approximately 500,000 people take the Common Test every year, and in order to be admitted to a university at the academic level of Institute of Science Tokyo, one would, by a simple calculation, need to be in roughly the top 2–3%. In addition, many students in Japan invest a considerable amount of time into university entrance examinations, so being in the top 2–3% is by no means a low level. To use an analogy, it is like trying to pick a fruit at the top of a Sequoia sempervirens. I want to pick a fruit from a lower tree.
On the other hand, in the admissions method that places emphasis on research projects, although it depends on the school to which one applies, about 40 people apply, around 12 pass the first-stage selection, and finally about 6 are admitted. The level of the research projects of successful applicants is very high for high-school students, but it is by no means an unreachable level, and I am exploring research-project topics that could reach a level capable of being admitted through this route. For that purpose, I want to organize my current thinking and hear opinions from outside. For what I mean by “level” in this post, please refer to this.
As a premise, although this is a selection process for the School of Computing, the range of topics that can be submitted as research projects is quite broad. One may write about mathematical investigations. One may write about what one learned through a programming class that one planned oneself. One may write about what one learned from competitions one participated in, such as competitive programming or the Chemistry Olympiad. It is also acceptable to write about how one’s experience studying abroad led one to think about something connected to the field of computing.
Also, if one passes the first-stage selection, there is an interview, where questions and answers are based on the contents one wrote. Therefore, even though one can use the Internet and AI during the research stage, if one writes something too far removed from one’s own actual ability, it will quickly be exposed. I also expect that such superficial thinking would be quickly seen through and would not even pass the first-stage selection.
There are also restrictions on the submission format of the research project: four pages in a Word file, font size 10 points or larger, with a free format. The research project may be carried out individually by the applicant, individually while receiving guidance from a teacher or others, or collaboratively.
At first, based on the theory of change andoriginal seeing, I conducted observations and analysis based on successful applicants’ accounts and examples published by the university. (LW readers are probably likely to know them, but I feel that those two essays are extremely valuable for producing results, so if you have not read them, I recommend doing so.)
What I learned from the observations and analysis
First, when the contents of the research projects submitted by successful applicants were reduced to a very rough scale that would be quite crude for rationalists but understandable to anyone, I found that they ranged from early undergraduate to senior undergraduate level.
I found that a certain degree of novelty in the research content is valued.
I found that the entire process of discovering a problem in the field oneself, thinking of a solution oneself, and verifying it is valued.
I found that the attitude of pursuing something one personally wondered about and actually experimenting with it is valued.
I found that it is valued when what one learned through the research is connected to what one wants to study at university in the future.
The research content itself does not necessarily have to be advanced. What is highly valued is an attempt to solve an unresolved problem that is practically problematic in that field, and the inclusion of experiences and insights specific to that research.
I found that it is highly valued when the contents of the report make it possible for the university to judge that the applicant fits the admissions policy issued by the university.
Finding a research project that satisfies many of these conditions is certainly difficult, but there seems to be a certain level that valued research reaches. Conversely, I thought that if I could meet that level, it would help differentiate me from other applicants and increase the probability of passing the first-stage selection.
However, there are also many problems.
First, the number of research projects on which the observations and analysis above are based is small. This admissions system itself is relatively new when viewed in the history of Japanese university entrance examinations, only about six people are admitted each year, and very few people record their experience for later applicants. In addition, because there is no data in my investigation about what kinds of projects were submitted by people who failed the first-stage selection, the overall level of applicants is unknown, and I expect that there is selection bias among successful applicants in terms of which people choose to publish their projects.
The rise of AI is also a problem. Applicants are expected to have at least a certain level of interest and ability in information-related fields, and I expect that they are using models at the 5.6 Sol level in some form, although the frequency of use may differ. Applicants can also be expected to have at least some skill in how to use them, so the level of research-project content may be raised substantially. In other words, I expect that rather than how advanced the content is, things that are difficult for AI to substitute for, such as experiments and statistics, will become differentiating factors. Therefore, the difficulty level of research examples produced by the university last year or several years ago may not be very useful for estimating the current competitive level.
Among the activities I have carried out so far, there is almost nothing that seems likely to reach a sufficient level as a research project. I think that I have accumulated enough insights about crystallized intelligence and fluid intelligence that I might be able to write a fairly interesting article in that area, but this has little primary-research element and has a large literature-review and synthesis aspect, which is exactly the kind of direction in which differentiation is difficult under evaluation standards in the AI era.
I have nobody I can consult with. In successful applicants’ accounts, they received feedback in some form from university students or researchers through social media such as X (Twitter). However, I do not use X in the first place, and I judged that it would be quite difficult in my current situation to meet someone I could consult about a topic like this, so I decided to post on LW.
I expect that AI will improve the quality and level of all applicants’ projects, but I do not intend to use AI for anything other than understanding knowledge and theory and proofreading writing. (See this famous article.) In my view, AI is the tool I know whose output usefulness depends most strongly on the context and amount of knowledge possessed by the user. For example, even if the same question is entered, I think readers of this Quick Take will intuitively understand that the usefulness of the answer will differ considerably between the context of someone who has read a large amount of discussion in the rationalist community and the context of someone who usually only uses AI for things like relationship advice. To put it further: if you do not know an answer, you can ask AI and find out. If you know what you need to investigate but do not know the terminology, you can ask AI and find out. But if you do not know that the domain or concept itself exists, you cannot even formulate the question to ask AI. In my own case, for example, it took me quite a long time to learn that the kind of discussions I was looking for were grouped under the name “rationality.” Especially from the Japanese-language sphere, there are very few paths that lead there. That is why I want to study systematically at university.
Research topics I am currently considering
Could fluid intelligence be substituted for to a considerable extent by crystallized intelligence? Fluid intelligence certainly seems to exist. However, from the perspective of producing results toward some terminal goal, could scaling meta-strategies and external tools eventually allow people to produce roughly the same level of results regardless of differences in individual fluid intelligence?
This may become an article separate from the research project because it seems likely to be useful to people, but: concrete methods for using spaced learning with Anki to efficiently increase crystallized intelligence in a specific context one wants, and thereby make judgments and think in a way that is pseudo-expert-like or ultimately close to that of an expert. (At least within what I searched, I could not find an LW article that directly dealt with this specific practical method.) This is based on the hypothesis that a considerable part of thinking is supported by words that label concepts. For those who are not very familiar with spaced learning, this is also a well-known article, and I recommend referring to it if you have not read it.
A thought experiment about creating an AGI specialized for the development of science. TsviBT writes the following in an article: “Current AI got capabilities through a path other than accumulating crystallized intelligence generated by its own fluid intelligence. That is, it got capabilities by, in a broad sense, ‘copying’ human crystallized intelligence that was originally generated by human fluid intelligence. This explanation seems able to account for most of the rapid increase in capability of current AI without invoking the idea that current AI has large fluid intelligence. In addition, given the apparent gap in which humans still remain far ahead in general fluid intelligence, I do not know how we could be very strongly confident that we already possess most or all of the ideas needed for AGI.” I think that human fluid intelligence and crystallized intelligence may ultimately have been formed through observations of the real world. Also, human working memory is limited, and the amount of time available to one human is limited, and there are few people who have simultaneously mastered distant fields such as mathematics and archaeology. Therefore, there may be unexplored transfers of methods between fields, such as meta-strategies that are common in archaeology but almost never used in mathematics. If AI were given an environment in which it could directly observe and experiment on the real world, might it therefore be possible to make new discoveries through cross-field combinations that humans have not sufficiently tried before?
AI’s output changes greatly for the same question depending on the context it is given, but can variables other than the prompt cause the model side to introduce contexts that the questioner themself does not know? If that is possible, then even if a user is not proficient in the field being asked about, could they obtain answers that are more useful for their objective?
Conclusion
These are still immature ideas, so there may be parts that I have not organized well. I would like people to point out the rough parts.
Anyone who has read this is welcome to point out or advise me about any part of it.
At present, I do not deeply understand each individual field, so if there is a problem currently being worked on in a field that a reader is knowledgeable about and that seems like it could make an interesting research project for a high-school student, I would like as many ideas as possible right now regardless of the field or level, so please tell me in the comments. (I am prepared to study for it.)
At least at the stage of generating research-topic ideas, I have not used AI. This is because in fields where I still possess only shallow context, I can only draw shallow answers from AI, and because if most applicants are using it in some form, there is a possibility that ideas will become homogenized in some way. What kinds of research topics would current AI be unlikely to suggest from a general prompt? Also, what kinds of research activities are difficult to substitute for even with AI? (For example, real-world experiments or surveys involving people.)
Is there a way to turn my current candidates, which lean toward thought experiments and literature reviews, into something closer to primary research?
Is there anyone who is interested in one of the research-topic candidates I am currently considering and would be willing to provide even a little consultation or advice?
In translating from Japanese into English, I used this prompt:
“Translate this Japanese directly into English without changing the content, theory, or body text. Do not add or remove information, and do not improve it through paraphrasing.”
Welcome! This is quite long for a shortform. Also, the topic is complex, and as you say most LessWrong readers are probably unfamiliar with the details of Japanese education system. But perhaps most importantly, it is not obvious what exactly is your opinion/proposal—could you please make a short summary? (And maybe report this as an article, with the short summary at the top.)
I took one of several ideas I had and developed it as concretely as I could at this point, then published it as a post. I also wrote a summary of the research, and someone left a helpful comment, so I’d appreciate it if you could take a look.
I will write my response here, not to derail the discussion under the article.
The introduction section is amazing! Maybe it’s because I am recently playing with an AI, but I can empathize with the “I can do this better”, “no wait, I can do this even better”, “wait again, there is actually even much better way to do this” sequence experienced on a scale of just several days.
Iterated improvement, but iterated so quickly that sometimes you just throw away version N even before it is finished, because updating to version N+1 and using that will just be so much faster. Yay!
(Also, beware, because somewhere along that way is AI psychosis waiting to get you.)
But also, specific examples help me understand your position better.
I suspect there will be a difference in results between people who just use an AI merely as a tool to run commands expressed in human language, and those who also use an AI to reflect on their own work.
I typically use the AI to write code, but I keep saying things like “does this make sense to you?” “what is your opinion on this?” “would you suggest an alternative approach?” “if this is somehow a not optimal or not standard approach, tell me” etc. Most of the time the AI just okays whatever I wrote, but once in a while it makes a valuable suggestion. Or, once I was not sure which parts of the project can the AI do reliably and which ones probably not… and then it occurred to me that I could actually ask the AI this very question. (Though maybe I should not blindly trust its answer.) Basically, AI is a tool that produces not only code and text, but also thoughts, and you should use it like that, too.
...sorry, this is unrelated to your topic, but the new version has inspired me to write something. ;)
Thank you for your reply. Your comment helped me put into words more clearly what I want to communicate to other people.
For example, even if I ask an AI for ways to improve something I am trying to accomplish with Python, or ask it to explore whether there is a better method, the AI can remain attached to the assumption that the terminal goal should be achieved using Python, even when extended search is included.
As a result, even after several rounds of improvement, and even if the AI evaluates the solution as “almost fully optimized,” there may actually be a fundamentally different method that does not use Python and works far better.
An even bigger problem is that, although the AI is supposedly thinking about methods for achieving the terminal goal, it may fail to propose such completely different options in the first place.
This suggests that the act of instructing the AI to “improve this” or “look for another method” may itself implicitly fix the current method and conversational context as assumptions, thereby restricting the search range. In other words, there may be a problem in which the very prompt used to ask the AI to explore ends up narrowing the search space.
That is frightening.
For every prompt I enter, even if the AI claims that it has “improved” something, is that really a major improvement?
At a microscopic level, a discovery may seem like a paradigm shift, but from a more macroscopic level, it may be only a very small improvement, and there may be a method that improves things far more. I suspect that such methods probably do exist. And proving that “no better method exists” is extremely difficult.
What is even more frightening is that this macroscopic improvement might, from the perspective of someone in a completely different field, be such a natural option that the AI would propose it immediately in the very first interaction.
If so, why should I be able to say, “This is the best solution,” about the result of continuing to search within a context that began from my own narrow knowledge?
Perhaps the real issue is not how intelligently the AI can improve a solution, but rather which possibilities for improvement are allowed to enter the search space, and which possibilities become invisible from the beginning because of the initial context it was given.
Another problem is that even a prompt that appears to broaden the search space, such as:
“Generate methods that three experts from different fields, who know nothing about the current solution, might propose if they were shown only the terminal goal.”
may itself unintentionally narrow the search space.
The moment I specify “three people,” “experts,” and “different fields,” I am already imposing a new framework on the directions the AI searches. Methods that only a non-expert might think of, methods that do not fit within existing disciplinary categories, or methods that question whether the terminal goal even needs to be achieved directly at all may instead be excluded from the search.
In other words, the very instructions added in order to broaden the search range may create a new search boundary.
Worse still, despite trying this many different prompts, in the Anki-efficiency example from my previous article, I have so far not found any proposal that goes beyond “use an AI agent.”
But looking back at the reasoning process up to this point, how can I be confident that no further paradigm-shift-level improvement exists?
Rather, the current search process itself may already be strongly constrained by a particular context or framework of thought. If that is the case, then no matter how many prompts I try, and even if the AI evaluates the result as “there is little room left for major improvement,” that would only be an evaluation within the search space that is currently visible.
The real problem is that, based only on the current reasoning process, I can hardly rule out the possibility that there are still undiscovered solutions located in completely different fields or at completely different levels of abstraction.
Do you think this is a serious problem? In my own mind, I had considered it important enough to be comparable to issues such as AI safety. However, given that the main article was classified as a Personal Blog post, I can think of several possibilities.
This research topic is already somewhat obvious on LessWrong, and I may in fact have been overestimating its importance.
Judging from the comments, I may not have articulated the seriousness of the problem well enough.
I may have focused too much on personal circumstances such as the admissions system, which made it harder for the issue to come across as a more general problem.
I had thought of LessWrong as a place where a wide range of important AI-related problems are discussed. However, in practice, it may be more focused on topics such as AI safety, and may not necessarily be the most appropriate place to discuss this kind of issue.
I would like to hear your candid opinion on this.
If you agree with this concern even to some extent, what would be a good way to gather a wider range of opinions on this topic from more people?
I think I see where you are coming from. The repeated experience that “this can be done much better” can turn into a constant suspicion that whatever you are doing is still a few such iterations away from optimal. The more we can go meta, the more impatient we become with not doing so.
I am also afraid that this feeling is related to AI psychosis. I consider myself psychologically stable, and yet recently the AI experience is quite addictive for me. It gives the feeling of unlimited possibilities. And the outcome is mixed: sometimes new possibilities open, sometimes it was all a hallucination.
Today I decided to take a one-day break from the AI, to breathe deeply and let my brain relax. Also, I think it is okay to go forward at not the maximum speed—if further improvements are possible, we will get there tomorrow. This is just an intuition, but I think that there is a cycle of: “invent an improvement, test the improvement, get experience, based on the experience propose a new improvements” and if we advance too fast, we don’t get time to do the “get experience” part properly.
And now that I think about it, this could be related to why AIs hallucinate so much. They actually “exist” only during the dialog, when we usually ask them to invent and build new things; they rarely get an opportunity to experience them. Not sure what such experience would even look like. For humans, it’s like “you have this new idea that you think is cool… keep applying it in random different situations, and then let’s see if you still approve of it afterwards, or if you came to some more nuanced understanding”. But the AI typically only applies its idea once, to the problem at hand. We usually do not even have random different situations for the AI to train on.
So maybe a better approach is “make a small progress, then stabilize your approach, and test the results”. In the spirit of “minimum viable product” that companies use, what is the “minimum viable improvement” you would propose? What is the smallest change that would clearly improve the status quo? Like in software we build version 1.0 and leave the other cool ideas for version 2.0 and later, what would be the version 1.0 of the improved education? Perhaps try to get this somehow implemented and tested first. Basically what I am trying to avoid is keeping the situation in constant motion, where everything is possible, but nothing is working out of the box. I mean, yeah, in long term, the constant motion is the source of progress. But along the way, we need some stable milestones. Or, to use a different analogy, mountain climbers sometimes make anchors, even if that slows them down, but it saves them when they lose the balance.
Also, smaller increments are easier to communicate. The best vision will be useless if no one actually tries it, and people need to understand it first.
To answer your question, I am not really afraid of not using our abilities to go meta to the fullest. The AIs are here, and they are going to stay. Everything we didn’t do today, we can still do tomorrow. Whatever limitations our attempts to generalize things meet.… when things calm down, when today’s “exciting new idea” becomes tomorrow “boring normal”, I think tomorrow we will notice the limitations and try to exceed them.
(This makes me suspect that I may have an instinctive fear that AIs are scarce and someone is going to take them away soon. A natural response to “we didn’t have them in the past, so it makes sense to see them instinctively as a rare thing”. But they will be here the next year, only more.)
So I would suggest to make a 1.0 proposal of the improvement, and try to communicate that. (Even if you already have versions 2.0 and 3.0 in your mind.)
Thank you for the truly useful point—one that I probably would not have noticed if I had consulted AI about it.
I came to understand firsthand what this blog had warned about. I thought I had been careful, but at some point I may have developed something close to AI psychosis. I do not usually ask AI for its opinions very much, but even when I was only trying to draw out knowledge, I may have been guiding it in a way that implied that this idea was excellent and useful.
Normally, I feel that I also have a way of thinking close to:
“Come up with an improvement, try the improvement, gain experience, and based on that experience, propose a new improvement.”
But this time, because of my desire to try new ideas with AI, I ended up disrupting the form that this feedback loop should originally have taken.
What I should have done was first try doing something, observe what happened, and then ask other people for their opinions together with those results. Otherwise, it is difficult for people to give concrete suggestions or impressions, and I feel that it was somewhat insincere of me to ask only for opinions without first taking that action.
This was also my first time posting on LW. Until now, I had only been reading articles, but I realized how valuable it is to actually have other people point things out to me.
Also, the phrase:
“a state where everything seems possible, but nothing works as-is, and the situation is constantly moving”
really resonated with me. For the past few days, I had had the feeling that I did not actually understand anything concrete about this idea, despite my expectations for it. Perhaps I had been refusing to observe the possibility that the idea I considered grand might not actually be as substantial as I expected.
However, I am still interested in this idea, so I plan to spend a few days making and thinking through an even smaller first step toward version 1.0, and then publish it.
“Smaller increments are also easier to communicate to other people. No matter how good the vision is, it is useless if nobody actually tries it. And for people to try it, they first need to understand it.”
These words of yours may become something I continue to value as an insight about communicating ideas to other people.
“If today’s ‘exciting new idea’ becomes tomorrow’s ‘boring obvious thing,’ then I think tomorrow’s version of us will notice its limitations and try to go beyond them.”
I especially strongly agree with this.
I also realized, more calmly, that as versions 1, 2, and 3 continue to develop, if this idea is genuinely useful and if the way I approach it is sincere in how I present it to other people, then perhaps people will naturally start giving their opinions.
I admit that my idea was not very clear. Once you pointed it out specifically, I realized that it was far too abstract for other people to properly evaluate. I was also relying too much on others, almost as if I were simply throwing the idea at an LLM and expecting it to work things out for me.
As you suggested, I’m going to think it through more carefully and develop the idea further before posting it as an article.
Purpose
I want to make better discoveries about a research-project topic to use for the comprehensive selection admissions process of the School of Computing at Institute of Science Tokyo (a Japanese university).
Background
Considering the distribution of Less Worng users, there do not seem to be very many Japanese users, so in order to explain the background, I need to give at least a minimal explanation of the Japanese university admissions system.
First, in order to apply to a Japanese national university, students first take something called the Common Test in January. Then, in February, they take an individual university examination (English, mathematics, chemistry, and physics). Some people may wonder why there are two examinations. The Common Test is administered by the national government, whereas the individual examinations are created by each university, and their difficulty and tendencies differ greatly between universities. Each university converts the Common Test score according to a certain weighting, and admission is decided based on the sum of that score and the score from the university’s individual examination.
However, each university also has distinctive admissions methods separate from this. For example, in the admissions method I am considering, the first-stage selection is based only on the Common Test and the evaluation of a research project conducted before the admissions process, and the second-stage selection is based only on an interview of about 30 minutes.
The advantage of this method is obvious. Compared with the former method, the required academic level is relatively lower. Under the ordinary method, approximately 500,000 people take the Common Test every year, and in order to be admitted to a university at the academic level of Institute of Science Tokyo, one would, by a simple calculation, need to be in roughly the top 2–3%. In addition, many students in Japan invest a considerable amount of time into university entrance examinations, so being in the top 2–3% is by no means a low level. To use an analogy, it is like trying to pick a fruit at the top of a Sequoia sempervirens. I want to pick a fruit from a lower tree.
On the other hand, in the admissions method that places emphasis on research projects, although it depends on the school to which one applies, about 40 people apply, around 12 pass the first-stage selection, and finally about 6 are admitted. The level of the research projects of successful applicants is very high for high-school students, but it is by no means an unreachable level, and I am exploring research-project topics that could reach a level capable of being admitted through this route. For that purpose, I want to organize my current thinking and hear opinions from outside. For what I mean by “level” in this post, please refer to this.
As a premise, although this is a selection process for the School of Computing, the range of topics that can be submitted as research projects is quite broad. One may write about mathematical investigations. One may write about what one learned through a programming class that one planned oneself. One may write about what one learned from competitions one participated in, such as competitive programming or the Chemistry Olympiad. It is also acceptable to write about how one’s experience studying abroad led one to think about something connected to the field of computing.
Also, if one passes the first-stage selection, there is an interview, where questions and answers are based on the contents one wrote. Therefore, even though one can use the Internet and AI during the research stage, if one writes something too far removed from one’s own actual ability, it will quickly be exposed. I also expect that such superficial thinking would be quickly seen through and would not even pass the first-stage selection.
There are also restrictions on the submission format of the research project: four pages in a Word file, font size 10 points or larger, with a free format. The research project may be carried out individually by the applicant, individually while receiving guidance from a teacher or others, or collaboratively.
At first, based on the theory of change and original seeing, I conducted observations and analysis based on successful applicants’ accounts and examples published by the university. (LW readers are probably likely to know them, but I feel that those two essays are extremely valuable for producing results, so if you have not read them, I recommend doing so.)
What I learned from the observations and analysis
First, when the contents of the research projects submitted by successful applicants were reduced to a very rough scale that would be quite crude for rationalists but understandable to anyone, I found that they ranged from early undergraduate to senior undergraduate level.
I found that a certain degree of novelty in the research content is valued.
I found that the entire process of discovering a problem in the field oneself, thinking of a solution oneself, and verifying it is valued.
I found that the attitude of pursuing something one personally wondered about and actually experimenting with it is valued.
I found that it is valued when what one learned through the research is connected to what one wants to study at university in the future.
The research content itself does not necessarily have to be advanced. What is highly valued is an attempt to solve an unresolved problem that is practically problematic in that field, and the inclusion of experiences and insights specific to that research.
I found that it is highly valued when the contents of the report make it possible for the university to judge that the applicant fits the admissions policy issued by the university.
Finding a research project that satisfies many of these conditions is certainly difficult, but there seems to be a certain level that valued research reaches. Conversely, I thought that if I could meet that level, it would help differentiate me from other applicants and increase the probability of passing the first-stage selection.
However, there are also many problems.
First, the number of research projects on which the observations and analysis above are based is small. This admissions system itself is relatively new when viewed in the history of Japanese university entrance examinations, only about six people are admitted each year, and very few people record their experience for later applicants. In addition, because there is no data in my investigation about what kinds of projects were submitted by people who failed the first-stage selection, the overall level of applicants is unknown, and I expect that there is selection bias among successful applicants in terms of which people choose to publish their projects.
The rise of AI is also a problem. Applicants are expected to have at least a certain level of interest and ability in information-related fields, and I expect that they are using models at the 5.6 Sol level in some form, although the frequency of use may differ. Applicants can also be expected to have at least some skill in how to use them, so the level of research-project content may be raised substantially. In other words, I expect that rather than how advanced the content is, things that are difficult for AI to substitute for, such as experiments and statistics, will become differentiating factors. Therefore, the difficulty level of research examples produced by the university last year or several years ago may not be very useful for estimating the current competitive level.
Among the activities I have carried out so far, there is almost nothing that seems likely to reach a sufficient level as a research project. I think that I have accumulated enough insights about crystallized intelligence and fluid intelligence that I might be able to write a fairly interesting article in that area, but this has little primary-research element and has a large literature-review and synthesis aspect, which is exactly the kind of direction in which differentiation is difficult under evaluation standards in the AI era.
I have nobody I can consult with. In successful applicants’ accounts, they received feedback in some form from university students or researchers through social media such as X (Twitter). However, I do not use X in the first place, and I judged that it would be quite difficult in my current situation to meet someone I could consult about a topic like this, so I decided to post on LW.
I expect that AI will improve the quality and level of all applicants’ projects, but I do not intend to use AI for anything other than understanding knowledge and theory and proofreading writing. (See this famous article.) In my view, AI is the tool I know whose output usefulness depends most strongly on the context and amount of knowledge possessed by the user. For example, even if the same question is entered, I think readers of this Quick Take will intuitively understand that the usefulness of the answer will differ considerably between the context of someone who has read a large amount of discussion in the rationalist community and the context of someone who usually only uses AI for things like relationship advice. To put it further: if you do not know an answer, you can ask AI and find out. If you know what you need to investigate but do not know the terminology, you can ask AI and find out. But if you do not know that the domain or concept itself exists, you cannot even formulate the question to ask AI. In my own case, for example, it took me quite a long time to learn that the kind of discussions I was looking for were grouped under the name “rationality.” Especially from the Japanese-language sphere, there are very few paths that lead there. That is why I want to study systematically at university.
Research topics I am currently considering
Could fluid intelligence be substituted for to a considerable extent by crystallized intelligence? Fluid intelligence certainly seems to exist. However, from the perspective of producing results toward some terminal goal, could scaling meta-strategies and external tools eventually allow people to produce roughly the same level of results regardless of differences in individual fluid intelligence?
This may become an article separate from the research project because it seems likely to be useful to people, but: concrete methods for using spaced learning with Anki to efficiently increase crystallized intelligence in a specific context one wants, and thereby make judgments and think in a way that is pseudo-expert-like or ultimately close to that of an expert. (At least within what I searched, I could not find an LW article that directly dealt with this specific practical method.) This is based on the hypothesis that a considerable part of thinking is supported by words that label concepts. For those who are not very familiar with spaced learning, this is also a well-known article, and I recommend referring to it if you have not read it.
A thought experiment about creating an AGI specialized for the development of science. TsviBT writes the following in an article: “Current AI got capabilities through a path other than accumulating crystallized intelligence generated by its own fluid intelligence. That is, it got capabilities by, in a broad sense, ‘copying’ human crystallized intelligence that was originally generated by human fluid intelligence. This explanation seems able to account for most of the rapid increase in capability of current AI without invoking the idea that current AI has large fluid intelligence. In addition, given the apparent gap in which humans still remain far ahead in general fluid intelligence, I do not know how we could be very strongly confident that we already possess most or all of the ideas needed for AGI.” I think that human fluid intelligence and crystallized intelligence may ultimately have been formed through observations of the real world. Also, human working memory is limited, and the amount of time available to one human is limited, and there are few people who have simultaneously mastered distant fields such as mathematics and archaeology. Therefore, there may be unexplored transfers of methods between fields, such as meta-strategies that are common in archaeology but almost never used in mathematics. If AI were given an environment in which it could directly observe and experiment on the real world, might it therefore be possible to make new discoveries through cross-field combinations that humans have not sufficiently tried before?
AI’s output changes greatly for the same question depending on the context it is given, but can variables other than the prompt cause the model side to introduce contexts that the questioner themself does not know? If that is possible, then even if a user is not proficient in the field being asked about, could they obtain answers that are more useful for their objective?
Conclusion
These are still immature ideas, so there may be parts that I have not organized well. I would like people to point out the rough parts.
Anyone who has read this is welcome to point out or advise me about any part of it.
At present, I do not deeply understand each individual field, so if there is a problem currently being worked on in a field that a reader is knowledgeable about and that seems like it could make an interesting research project for a high-school student, I would like as many ideas as possible right now regardless of the field or level, so please tell me in the comments. (I am prepared to study for it.)
At least at the stage of generating research-topic ideas, I have not used AI. This is because in fields where I still possess only shallow context, I can only draw shallow answers from AI, and because if most applicants are using it in some form, there is a possibility that ideas will become homogenized in some way. What kinds of research topics would current AI be unlikely to suggest from a general prompt? Also, what kinds of research activities are difficult to substitute for even with AI? (For example, real-world experiments or surveys involving people.)
Is there a way to turn my current candidates, which lean toward thought experiments and literature reviews, into something closer to primary research?
Is there anyone who is interested in one of the research-topic candidates I am currently considering and would be willing to provide even a little consultation or advice?
In translating from Japanese into English, I used this prompt:
“Translate this Japanese directly into English without changing the content, theory, or body text. Do not add or remove information, and do not improve it through paraphrasing.”
Welcome! This is quite long for a shortform. Also, the topic is complex, and as you say most LessWrong readers are probably unfamiliar with the details of Japanese education system. But perhaps most importantly, it is not obvious what exactly is your opinion/proposal—could you please make a short summary? (And maybe report this as an article, with the short summary at the top.)
I took one of several ideas I had and developed it as concretely as I could at this point, then published it as a post. I also wrote a summary of the research, and someone left a helpful comment, so I’d appreciate it if you could take a look.
https://www.lesswrong.com/posts/ozqsjiijrGAyx6spr/would-this-research-outline-be-interesting-i-would-like-to
Nice!
I will write my response here, not to derail the discussion under the article.
The introduction section is amazing! Maybe it’s because I am recently playing with an AI, but I can empathize with the “I can do this better”, “no wait, I can do this even better”, “wait again, there is actually even much better way to do this” sequence experienced on a scale of just several days.
Iterated improvement, but iterated so quickly that sometimes you just throw away version N even before it is finished, because updating to version N+1 and using that will just be so much faster. Yay!
(Also, beware, because somewhere along that way is AI psychosis waiting to get you.)
But also, specific examples help me understand your position better.
I suspect there will be a difference in results between people who just use an AI merely as a tool to run commands expressed in human language, and those who also use an AI to reflect on their own work.
I typically use the AI to write code, but I keep saying things like “does this make sense to you?” “what is your opinion on this?” “would you suggest an alternative approach?” “if this is somehow a not optimal or not standard approach, tell me” etc. Most of the time the AI just okays whatever I wrote, but once in a while it makes a valuable suggestion. Or, once I was not sure which parts of the project can the AI do reliably and which ones probably not… and then it occurred to me that I could actually ask the AI this very question. (Though maybe I should not blindly trust its answer.) Basically, AI is a tool that produces not only code and text, but also thoughts, and you should use it like that, too.
...sorry, this is unrelated to your topic, but the new version has inspired me to write something. ;)
Thank you for your reply. Your comment helped me put into words more clearly what I want to communicate to other people.
For example, even if I ask an AI for ways to improve something I am trying to accomplish with Python, or ask it to explore whether there is a better method, the AI can remain attached to the assumption that the terminal goal should be achieved using Python, even when extended search is included.
As a result, even after several rounds of improvement, and even if the AI evaluates the solution as “almost fully optimized,” there may actually be a fundamentally different method that does not use Python and works far better.
An even bigger problem is that, although the AI is supposedly thinking about methods for achieving the terminal goal, it may fail to propose such completely different options in the first place.
This suggests that the act of instructing the AI to “improve this” or “look for another method” may itself implicitly fix the current method and conversational context as assumptions, thereby restricting the search range. In other words, there may be a problem in which the very prompt used to ask the AI to explore ends up narrowing the search space.
That is frightening.
For every prompt I enter, even if the AI claims that it has “improved” something, is that really a major improvement?
At a microscopic level, a discovery may seem like a paradigm shift, but from a more macroscopic level, it may be only a very small improvement, and there may be a method that improves things far more. I suspect that such methods probably do exist. And proving that “no better method exists” is extremely difficult.
What is even more frightening is that this macroscopic improvement might, from the perspective of someone in a completely different field, be such a natural option that the AI would propose it immediately in the very first interaction.
If so, why should I be able to say, “This is the best solution,” about the result of continuing to search within a context that began from my own narrow knowledge?
Perhaps the real issue is not how intelligently the AI can improve a solution, but rather which possibilities for improvement are allowed to enter the search space, and which possibilities become invisible from the beginning because of the initial context it was given.
Another problem is that even a prompt that appears to broaden the search space, such as:
“Generate methods that three experts from different fields, who know nothing about the current solution, might propose if they were shown only the terminal goal.”
may itself unintentionally narrow the search space.
The moment I specify “three people,” “experts,” and “different fields,” I am already imposing a new framework on the directions the AI searches. Methods that only a non-expert might think of, methods that do not fit within existing disciplinary categories, or methods that question whether the terminal goal even needs to be achieved directly at all may instead be excluded from the search.
In other words, the very instructions added in order to broaden the search range may create a new search boundary.
Worse still, despite trying this many different prompts, in the Anki-efficiency example from my previous article, I have so far not found any proposal that goes beyond “use an AI agent.”
But looking back at the reasoning process up to this point, how can I be confident that no further paradigm-shift-level improvement exists?
Rather, the current search process itself may already be strongly constrained by a particular context or framework of thought. If that is the case, then no matter how many prompts I try, and even if the AI evaluates the result as “there is little room left for major improvement,” that would only be an evaluation within the search space that is currently visible.
The real problem is that, based only on the current reasoning process, I can hardly rule out the possibility that there are still undiscovered solutions located in completely different fields or at completely different levels of abstraction.
Do you think this is a serious problem? In my own mind, I had considered it important enough to be comparable to issues such as AI safety. However, given that the main article was classified as a Personal Blog post, I can think of several possibilities.
This research topic is already somewhat obvious on LessWrong, and I may in fact have been overestimating its importance.
Judging from the comments, I may not have articulated the seriousness of the problem well enough.
I may have focused too much on personal circumstances such as the admissions system, which made it harder for the issue to come across as a more general problem.
I had thought of LessWrong as a place where a wide range of important AI-related problems are discussed. However, in practice, it may be more focused on topics such as AI safety, and may not necessarily be the most appropriate place to discuss this kind of issue.
I would like to hear your candid opinion on this.
If you agree with this concern even to some extent, what would be a good way to gather a wider range of opinions on this topic from more people?
I think I see where you are coming from. The repeated experience that “this can be done much better” can turn into a constant suspicion that whatever you are doing is still a few such iterations away from optimal. The more we can go meta, the more impatient we become with not doing so.
I am also afraid that this feeling is related to AI psychosis. I consider myself psychologically stable, and yet recently the AI experience is quite addictive for me. It gives the feeling of unlimited possibilities. And the outcome is mixed: sometimes new possibilities open, sometimes it was all a hallucination.
Today I decided to take a one-day break from the AI, to breathe deeply and let my brain relax. Also, I think it is okay to go forward at not the maximum speed—if further improvements are possible, we will get there tomorrow. This is just an intuition, but I think that there is a cycle of: “invent an improvement, test the improvement, get experience, based on the experience propose a new improvements” and if we advance too fast, we don’t get time to do the “get experience” part properly.
And now that I think about it, this could be related to why AIs hallucinate so much. They actually “exist” only during the dialog, when we usually ask them to invent and build new things; they rarely get an opportunity to experience them. Not sure what such experience would even look like. For humans, it’s like “you have this new idea that you think is cool… keep applying it in random different situations, and then let’s see if you still approve of it afterwards, or if you came to some more nuanced understanding”. But the AI typically only applies its idea once, to the problem at hand. We usually do not even have random different situations for the AI to train on.
So maybe a better approach is “make a small progress, then stabilize your approach, and test the results”. In the spirit of “minimum viable product” that companies use, what is the “minimum viable improvement” you would propose? What is the smallest change that would clearly improve the status quo? Like in software we build version 1.0 and leave the other cool ideas for version 2.0 and later, what would be the version 1.0 of the improved education? Perhaps try to get this somehow implemented and tested first. Basically what I am trying to avoid is keeping the situation in constant motion, where everything is possible, but nothing is working out of the box. I mean, yeah, in long term, the constant motion is the source of progress. But along the way, we need some stable milestones. Or, to use a different analogy, mountain climbers sometimes make anchors, even if that slows them down, but it saves them when they lose the balance.
Also, smaller increments are easier to communicate. The best vision will be useless if no one actually tries it, and people need to understand it first.
To answer your question, I am not really afraid of not using our abilities to go meta to the fullest. The AIs are here, and they are going to stay. Everything we didn’t do today, we can still do tomorrow. Whatever limitations our attempts to generalize things meet.… when things calm down, when today’s “exciting new idea” becomes tomorrow “boring normal”, I think tomorrow we will notice the limitations and try to exceed them.
(This makes me suspect that I may have an instinctive fear that AIs are scarce and someone is going to take them away soon. A natural response to “we didn’t have them in the past, so it makes sense to see them instinctively as a rare thing”. But they will be here the next year, only more.)
So I would suggest to make a 1.0 proposal of the improvement, and try to communicate that. (Even if you already have versions 2.0 and 3.0 in your mind.)
Thank you for the truly useful point—one that I probably would not have noticed if I had consulted AI about it.
I came to understand firsthand what this blog had warned about. I thought I had been careful, but at some point I may have developed something close to AI psychosis. I do not usually ask AI for its opinions very much, but even when I was only trying to draw out knowledge, I may have been guiding it in a way that implied that this idea was excellent and useful.
Normally, I feel that I also have a way of thinking close to:
“Come up with an improvement, try the improvement, gain experience, and based on that experience, propose a new improvement.”
But this time, because of my desire to try new ideas with AI, I ended up disrupting the form that this feedback loop should originally have taken.
What I should have done was first try doing something, observe what happened, and then ask other people for their opinions together with those results. Otherwise, it is difficult for people to give concrete suggestions or impressions, and I feel that it was somewhat insincere of me to ask only for opinions without first taking that action.
This was also my first time posting on LW. Until now, I had only been reading articles, but I realized how valuable it is to actually have other people point things out to me.
Also, the phrase:
“a state where everything seems possible, but nothing works as-is, and the situation is constantly moving”
really resonated with me. For the past few days, I had had the feeling that I did not actually understand anything concrete about this idea, despite my expectations for it. Perhaps I had been refusing to observe the possibility that the idea I considered grand might not actually be as substantial as I expected.
However, I am still interested in this idea, so I plan to spend a few days making and thinking through an even smaller first step toward version 1.0, and then publish it.
“Smaller increments are also easier to communicate to other people. No matter how good the vision is, it is useless if nobody actually tries it. And for people to try it, they first need to understand it.”
These words of yours may become something I continue to value as an insight about communicating ideas to other people.
“If today’s ‘exciting new idea’ becomes tomorrow’s ‘boring obvious thing,’ then I think tomorrow’s version of us will notice its limitations and try to go beyond them.”
I especially strongly agree with this.
I also realized, more calmly, that as versions 1, 2, and 3 continue to develop, if this idea is genuinely useful and if the way I approach it is sincere in how I present it to other people, then perhaps people will naturally start giving their opinions.
Thank you for the advice.
Thank you for the advice.
I admit that my idea was not very clear. Once you pointed it out specifically, I realized that it was far too abstract for other people to properly evaluate. I was also relying too much on others, almost as if I were simply throwing the idea at an LLM and expecting it to work things out for me.
As you suggested, I’m going to think it through more carefully and develop the idea further before posting it as an article.