1stuserhere

Karma: 76

Hi, I’m Kunvar. I enjoy science and engineering. I’m currently studying the structures learned by trained neural networks and trying to understand what algorithms are implemented by these networks for various tasks.

You can find my various social profiles here, my personal website here, and reach out to me at kunvar@mechinterp.com

1stuserhere Oct 22, 2024, 9:13 PM
1 point
0
in reply to: leogao’s comment on: A Rocket–Interpretability Analogy

on the one hand, mechanistic understanding has historically underperformed as a research strategy,

Are you talking about ML or in general? What are you deriving this from?

1stuserhere Oct 2, 2024, 4:12 PM
8 points
7
on: “Slow” takeoff is a terrible term for “maybe even faster takeoff, actually”

I think it’s even more actively confusing because “smooth/continuous” takeoff not only could be faster in calendar time

We’re talking about two different things here: take-off velocity, and timelines. All 4 possibilities are on the table—slow takeoff/long timelines, fast takeoff/long timelines, slow takeoff/short timelines, fast takeoff/short timelines.

A smooth takeoff might actually take longer in calendar time if incremental progress doesn’t lead to exponential gains until later stages.

Honestly I’m surprised people are conflating timelines and takeoff speeds.

1stuserhere Aug 30, 2024, 7:09 PM
4 points
2
in reply to: Noa Nabeshima’s comment on: Are the majority of your ancestors farmers or non-farmers?
That’s interesting. On the recent episode of Dwarkesh Podcast with David Reich, at 1:18:00, there’s a discussion I’ll quote here:

There was a super interesting series of papers. They made many things clear but one of them was that actually the proportion of non-Africans ancestors who are Neanderthals is not 2%.

That’s the proportion of their DNA in our genomes today if you’re a non-African person. It’s more like 10-20% of your ancestors are Neanderthals. What actually happened was that when Neanderthals and modern humans met and mixed, the Neanderthal DNA was not as biologically fit.

The reason was that Neanderthals had lived in small populations for about half a million years since separating from modern humans—who had lived in larger populations—and had accumulated a large number, thousands of slightly bad mutations. In the mixed populations, there was selection to remove the Neanderthal ancestry. That would have happened very, very rapidly after the mixture process.

There’s now overwhelming evidence that that must have happened. If you actually count your ancestors, if you’re of non-African descent, how many of them were Neanderthals say, 70,000 years ago, it’s not going to be 2%. It’s going to be 10-20%, which is a lot.

Now I don’t know which paper this is referring to but it’s interesting nonetheless.

1stuserhere Aug 23, 2024, 6:06 PM
3 points
0
on: what becoming more secure did for me
Good post, thanks for sharing! found it somewhat relatable to my prior life experiences too

1stuserhere Aug 5, 2024, 7:52 PM
2 points
0
on: The ‘strong’ feature hypothesis could be wrong
Great essay!

I found it to be well written and articulates many of my own arguments in casual conversations well. I’ll write up a longer comment with some points I found interesting and concrete questions accompanying them sometime later.

1stuserhere Aug 5, 2024, 5:39 AM
2 points
0
on: Can We Predict Persuasiveness Better Than Anthropic?

For each of the resulting 1313 arguments, crowdworkers were first asked to rate their support of the corresponding claim on a Likert scale from 1 (“Strongly Oppose”) to 7 (“Strongly Oppose”).

You probably mean strongly oppose to strongly support

1stuserhere Jul 17, 2024, 4:30 PM
1 point
0
in reply to: Daniel Tan’s comment on: Daniel Tan’s Shortform
You’ll enjoy reading What Causes Polysemanticity? An Alternative Origin Story of Mixed Selectivity from Incidental Causes (link to the paper)

Using a combination of theory and experiments, we show that incidental polysemanticity can arise due to multiple reasons including regularization and neural noise; this incidental polysemanticity occurs because random initialization can, by chance alone, initially assign multiple features to the same neuron, and the training dynamics then strengthen such overlap.

1stuserhere Jul 17, 2024, 4:18 PM
3 points
0
in reply to: Daniel Tan’s comment on: Daniel Tan’s Shortform

If we train several SAEs from scratch on the same set of model activations, are they “equivalent”?

For SAEs of different sizes, for most layers, the smaller SAE does contain very high similarity with some of the larger SAE features, but it’s not always true. I’m working on an upcoming post on this.

1stuserhere Apr 18, 2024, 10:31 AM
4 points
0
on: How to accelerate recovery from sleep debt with biohacking?
This is purely anecdotal—supplementing sleep debt with cardio-intensive exercise works for me. For example, I usually need 7 hrs of sleep. If I sleep for only 5 hrs, I’m likely to feel a drop in mental sharpness around midway the next day. However, if I go for an hour long run, I miss that drop almost completely and feel just as good I normally would’ve with a complete sleep.

1stuserhere Nov 15, 2023, 12:42 PM
2 points
0
in reply to: Charlie Steiner’s comment on: A framing for interpretability
It’s also worth noting that LLMs are not learning directly from the raw input stream but from a crux of that data (LLMs learn on compressed data) i.e. the LLMs are fed tokenized data, and the tokenizers act as compressors. This benefits the models by enabling them to have a more information-rich context.

Mechanistic Interpretability Reading group

1stuserhere and woog

Sep 26, 2023, 4:26 PM

15 points

0 comments1 min readLW link

1stuserhere Sep 15, 2023, 11:38 AM
1 point
0
in reply to: StellaAthena’s comment on: Are Mixture-of-Experts Transformers More Interpretable Than Dense Transformers?
I think that the answer is no
In this “VRAM-constrained regime,” MoE models (trained from scratch) are nowhere near competitive with dense LLMs.
Curious whether your high-level thoughts on these topics still hold or have changed.

1stuserhere Jun 5, 2023, 9:30 PM
1 point
0
on: How to Think About Activation Patching
On a more narrow distribution this head could easily exhibit just one behaviour and eg seem like a monosemantic inductin head
induction* head

1stuserhere Apr 23, 2023, 12:59 PM
8 points
0
in reply to: Dan H’s comment on: What 2026 looks like (Daniel’s Median Future)
The 2023 predictions seem to hold up really well, so far, especially the SDM in interactive environment one, image synthesis, passing the bar exam, legal NLP systems, enthusiasm of programmers, and Elon Musk re-entering the space of building AI systems.

1stuserhere Feb 27, 2023, 4:50 PM
5 points
2
in reply to: Mary Chernyshenko’s comment on: How to Read Papers Efficiently: Fast-then-Slow Three pass method
Interesting perspective especially your comments on citations. Agreed with the diagrams/figures/tables being some of the most interesting parts of the paper, but I also try to find the problem that motivated the authors (which is frequently embedded better in the introduction imo than the abstract).

How to Read Papers Efficiently: Fast-then-Slow Three pass method

the gears to ascension, 1stuserhere and lastuserhere

Feb 25, 2023, 2:56 AM

36 points

4 comments4 min readLW link

(ccr.sigcomm.org)

1stuserhere Feb 23, 2023, 1:46 PM
3 points
1
in reply to: Aprillion’s comment on: AI alignment researchers don’t (seem to) stack
In this analogy, the trouble is, we do not know whether we’re building tunnels in parallel (same direction) or the opposite, or zig zag. The reason for that is a lack of clarity about what will turn out to be a fundamentally important approach towards building a safe AGI. So, it seems to me that for now, exploration for different approaches might be a good thing and the next generation of researchers does less digging and is able to stack more on the existing work

1stuserhere Feb 23, 2023, 1:41 PM
4 points
1
in reply to: Zach Stein-Perlman’s comment on: AI alignment researchers don’t (seem to) stack
I agree. It seems like striking a balance between exploration and exploitation. We’re barely entering the 2nd generation of alignment researchers. It’s important to generate new directions of approaching the problem especially at this stage, so that we have a better chance of covering more of the space of possible solutions before deciding to go in deeper. The barrier to entry also remains slightly lower in this case for new researchers. When some research directions “outcompete” other directions, we’ll naturally see more interest in those promising directions and subsequently more exploitation, and researchers will be stacking.

1stuserhere

Mechanis­tic In­ter­pretabil­ity Read­ing group

How to Read Papers Effi­ciently: Fast-then-Slow Three pass method

Mechanistic Interpretability Reading group

How to Read Papers Efficiently: Fast-then-Slow Three pass method