The Synthetic Ledger
How a generative video startup proved exactly what their model was eating.
This story sits inside our mission to make digital provenance effortless, universal, and trustworthy for the people and systems who depend on it. See Mission & Vision
Problem … When the studio says: "What's really in this model?"
FluxFrame builds a generative video model. Their training mix: licensed clips from stock providers and studios, fully synthetic renders from 3D pipelines, public domain footage.
A major studio is interested in partnering… but asks: "Can you prove you didn't quietly train on our shows or unlabeled user uploads?" "What percentage of your training data is synthetic?"
Internally:
- Some data is tagged "licensed" in a spreadsheet.
- Synthetic content is scattered.
- Public domain sources aren't clearly documented.
- Retracing exact training datasets for each model version is painful.
Summary:
- No single view of what went into which model.
- No precise synthetic vs non-synthetic breakdown.
- No trustable evidence for skeptical legal teams.
Solution … Dataset passports with synthetic ratios
Step 1: Register media datasets
Each dataset (or batch of clips) is registered with a dataset passport that includes: source type: "licensed", "synthetic", "public_domain", "user_generated", licensing terms and provider, jurisdiction / sovereign zone, content type (video, frames, audio), synthetic flag + generation pipeline metadata where relevant. Each passport is hashed, gets external anchors (e.g. OTS / IPFS), stores aggregate stats (e.g. duration, count, synthetic ratio).
Step 2: Build training sets as compositions of passports
When FluxFrame builds a training set for a model, they compose multiple dataset passports. Sovereign creates a composite training set passport with list of constituent datasets, percentage breakdown: licensed vs public vs synthetic, provider breakdown (e.g. StockHouse 40%, Studio Y 20%, Synthetic 40%).
Model-level provenance … The training diet becomes part of the model's identity
When a training run completes, FluxFrame logs composite training set passport ID(s), architecture and run metadata, version number / commit hash. Sovereign ties the model passport to the synthetic ratio of its training corpus, which licensed providers are involved, which public domain sources are included, and anchors the model passport externally.
Flow … What happens when a studio does diligence
Step 1: Request
A studio asks: "What's in this model's training set? Show us breakdowns and sources."
Step 2: Export
FluxFrame exports a media training ledger bundle: model passport, composite training set passport, underlying dataset passports relevant to the partner's questions. The bundle makes clear: total training hours, proportions: synthetic vs licensed vs public domain, which licensed providers are in the mix, which sovereign zones govern data residency.
Step 3: Review
Studio legal teams review the bundle: they see no unlabeled scraped streaming content, only licensed and declared sources.
Step 4: Partnership
The partnership moves forward with clear contractual references to the specific model and dataset passports.
"Instead of vague assurances, FluxFrame sends a cryptographically anchored training ledger."
Outcomes … Proof instead of promises
- FluxFrame secures deals with IP owners who want clarity, differentiates from competitors who say "trust us" without evidence.
- For studios: they can audit model diets, require specific dataset exclusions in future versions.
"We stopped hand-waving about our data mix. Now we show the ledger."
Product tie-in … Why this is SOVEREIGN\\PROVENANCE for media
Not just logs. Not just a data lake catalog. Sovereign treats each dataset and training run as a provenance artifact, encodes synthetic vs non-synthetic signals, licenses, jurisdictions, training composition. That becomes the basis for contract language, due diligence, risk assessment.
If your model touches media IP, you need a synthetic ledger.
Show partners what your model really trained on.
SOVEREIGN\PROVENANCE turns your training mix into a ledger the legal team can read.
Why this matters now
Cases like this are emerging because AI, synthetic media, and global information flows have made provenance a necessity rather than a luxury. When it's no longer obvious who made what, when, or why, stories like this become the norm.
The Genesis Moment explains the broader context for this kind of provenance work. Read Genesis.
Work like this, when implemented on the SOVEREIGN\PROVENANCE platform, is governed by The Covenant of Restraint, our ethical framework for how we build and deploy provenance systems.
Governed by The Covenant of Restraint, aligned with our Mission & Vision, and grounded in the Genesis Moment.