Thinking on this more, I’m curious how/whether copyright is considered in Plan A’s Total Research Transparency (a proposal I broadly like.) Rereading the default proposal for TRT, it calls to restrict most training data:
AI model weights, significant fractions of the training data, and a small amount of other sensitive information is prevented from leaving the datacenters.
I’m a bit surprised because I assumed “data” was part of “research”; most calls for “open science” include calls to publish the underlying data. And I assumed that if the goal of Plan A is to make it so that model progress is strictly gatekept on compute, in which case making training data also transparent/public would aid that goal.
And now I’m also curious what kind of training data would be published as part of TRT.
Thinking on this more, I’m curious how/whether copyright is considered in Plan A’s Total Research Transparency (a proposal I broadly like.) Rereading the default proposal for TRT, it calls to restrict most training data:
I’m a bit surprised because I assumed “data” was part of “research”; most calls for “open science” include calls to publish the underlying data. And I assumed that if the goal of Plan A is to make it so that model progress is strictly gatekept on compute, in which case making training data also transparent/public would aid that goal.
And now I’m also curious what kind of training data would be published as part of TRT.
Tagging @Thomas Larsen?