OpenAI won't say whose content trained its video tool. We found some clues.
A recent investigation explores the undisclosed training data used for OpenAI's video generation tool, Sora. Despite OpenAI's lack of transparency regarding the origins of its training content, researchers and journalists are actively seeking clues to identify potential sources. The inquiry aims to understand the dataset composition, which is critical for evaluating the model's biases, intellectual property implications, and the ethical considerations surrounding generative AI development. Early findings suggest a diverse range of content might have been utilized, though specific creators or copyright holders remain unconfirmed. This ongoing effort highlights the growing demand for greater accountability and transparency from AI developers regarding their data practices, especially as advanced generative models become more prevalent and influential across various industries. The investigation underscores the challenges in tracing digital content back to its original creators when it's incorporated into large-scale AI training datasets, prompting broader discussions on data provenance in the AI era.