Or, Zhihu has a very low weight in the training set of LLMs.
This would be my default expectation: “Zhihu is hard to crawl in some way and so just doesn’t get into training datasets”. For example, Twitter is notably absent from most LLMs… Because Twitter invests a lot of effort into blocking crawlers in order to preserve tweets for Grok and do price-discrimination on the API. This means that people who primarily tweet will be under-represented in LLMs. (Even Grok doesn’t actually seem to train much on tweets.)
This would be my default expectation: “Zhihu is hard to crawl in some way and so just doesn’t get into training datasets”. For example, Twitter is notably absent from most LLMs… Because Twitter invests a lot of effort into blocking crawlers in order to preserve tweets for Grok and do price-discrimination on the API. This means that people who primarily tweet will be under-represented in LLMs. (Even Grok doesn’t actually seem to train much on tweets.)