Very interesting post! Is there a plan to open-source this? This can be a very useful starting point for further research.
I am particularly interested in the results with bolding since they seem clean & counterintuitive. My quick replication with Claude suggests that de-bolding the right data (but not just the coding data as you did) does help. Claude also cannot replicate the initial 29-96 gap (base comes out 69% for me with chat template), potentially due to the question distribution being different. These replications could be off in subtle ways so open-sourcing would be really helpful!
Hi @Ziqian Zhong—sorry for the late reply! Unfortunately this is quite work from a few months ago with to plans to open source at the moment—however, I am happy to share with you the exact datasets/experiment details offline I used for the experiments, if that is useful—DM me!
Very interesting post! Is there a plan to open-source this? This can be a very useful starting point for further research.
I am particularly interested in the results with bolding since they seem clean & counterintuitive. My quick replication with Claude suggests that de-bolding the right data (but not just the coding data as you did) does help. Claude also cannot replicate the initial 29-96 gap (base comes out 69% for me with chat template), potentially due to the question distribution being different. These replications could be off in subtle ways so open-sourcing would be really helpful!
Hi @Ziqian Zhong—sorry for the late reply! Unfortunately this is quite work from a few months ago with to plans to open source at the moment—however, I am happy to share with you the exact datasets/experiment details offline I used for the experiments, if that is useful—DM me!