October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
VGSources
AI

OpenAI Appears to Have Trained Sora on Game Content

Sora has generated recognizable game-like footage, but those outputs cannot identify a specific training video or prove that a game publisher supplied it.

By VGSources Team 4 min read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Probably in some form, but the public evidence does not prove which game videos Sora was trained on—or that any particular publisher supplied them. OpenAI says Sora used publicly available, partnership, and internally developed data. Tests reported by TechCrunch and The Washington Post found game-like scenes, recognizable logos, and gameplay-style clips in Sora’s outputs. That supports an inference that game-related visual material influenced the model; it is not an itemized record of its training set.

What OpenAI has disclosed about Sora’s training data

OpenAI’s 2025 Sora System Card describes a mixture of publicly available data, proprietary data accessed through partnerships, and custom datasets developed in-house. It says publicly available data included machine-learning datasets and web crawls, and also refers to partnership data and human feedback. OpenAI names Shutterstock and Pond5 as examples of partners.

This disclosure identifies broad categories, not individual video files, game titles, publishers, uploaders, or the licensing status of each item. OpenAI’s general training explainer, updated in 2026, similarly describes foundation-model training as using publicly available internet information, third-party partner data, and information supplied or generated by users, trainers, and researchers. It does not provide an itemized Sora video inventory.

The Associated Press reported on February 15, 2024, that OpenAI had not disclosed the imagery and video sources used to train Sora. OpenAI later described data categories, but that broader disclosure still does not establish which game footage, if any, was included.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why Sora can generate game-like clips

TechCrunch reported that prompts including “Italian plumber game” produced game-like imagery, and said game content may have found its way into Sora’s training data. The Washington Post later reported that Sora could generate clips resembling Minecraft, game logos, and a streamer playing Civilization. Researchers quoted by the Post said such results suggested versions of originals may have appeared in training data, while cautioning that resemblance alone cannot establish direct copying from a rights holder.

Generative models learn visual patterns from examples. A model may produce familiar game aesthetics, interface elements, characters, or streamer setups because it encountered related material, without reproducing one complete source video verbatim. But the outputs do not reveal whether any influence came from an authorized partner file, a public webpage, a user upload, or another source. Public availability is not the same as permission, and a familiar-looking result cannot identify its chain of custody.

TechCrunch quoted intellectual-property attorney Joshua Weigensberg saying, “Training a generative AI model generally involves copying the training data.” That observation concerns the technical process; it does not identify which particular files Sora used or resolve whether copying in a given case was licensed or legally permitted. As Joanna Materzynska told The Washington Post, “The model is mimicking the training data. There’s no magic.” The important evidentiary limit is that a model’s output is not, by itself, a source log.

What game-like output does—and does not—show

Possible explanation What the reported evidence supports What remains unestablished
Exposure to game-related visual material Reported tests found game-like imagery, logos, and gameplay-style scenes in Sora outputs (TechCrunch, 2024; The Washington Post, 2025). The exact source files, titles, uploaders, and number of examples in training are not stated by those reports.
Direct use of a specific publisher’s archive The reported resemblance can motivate that question. No reviewed source identifies a specific game publisher’s archive as a Sora training source or establishes direct copying from a rights holder.
Influence from public or user-uploaded material OpenAI says its models use publicly available data and information supplied or generated by users, trainers, and researchers (OpenAI, 2025 and 2026). Those categories do not show whether a particular game clip was included, who uploaded it, or whether it was authorized.

The distinction matters: evidence that outputs resemble games is evidence about model behavior, not a complete provenance trail. It cannot establish that Nintendo, Microsoft, Mojang, Twitch, or another named rights holder provided footage.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does a reproduced game logo prove copyright infringement?

No. A recognizable logo or game-like scene does not, by itself, prove infringement. A legal assessment would need to examine what material was copied or used, how it was obtained, what the output reproduces, and the applicable law and defenses in the relevant jurisdiction. Those questions may differ by country and by the facts of a particular case; the reported tests do not answer them.

Training-data questions and output questions are related but distinct. Evidence that a model was exposed to material does not alone establish that a particular generated result is infringing. Conversely, an output that does not reproduce a source video exactly would not, on its own, settle whether the training process was lawful. The available reporting does not identify a source recording or resolve those legal issues.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What provenance tools can tell you

C2PA and other content-provenance mechanisms can help attach information about a digital asset’s origin or editing history. They may help label a generated output, where supported and preserved, but they do not establish which videos were used to train the model. Output provenance and training-data provenance are different records; one cannot substitute for the other.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from Patch Notes

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.