The Unpoisoned Lab
How a collaborative ML lab made every dataset contribution traceable… and reversible.
This story sits inside our mission to make digital provenance effortless, universal, and trustworthy for the people and systems who depend on it. See Mission & Vision
Problem … When "open contributions" become an attack surface
OpenLens Lab fine-tunes open models with open datasets, contributed curated subsets, annotation campaigns. The good news: they move fast and experiment widely. The bad news: they can't always tell who contributed what, what exactly was in each dataset version, which contributions went into which training runs.
A worrying event: The community flags harmful model behavior. The lab suspects that a malicious dataset contributed to a recent fine-tune. They don't have a clear way to trace it.
Summary:
- No signed, per-contribution provenance.
- Weak linkage between dataset versions and training runs.
- Hard to remove "just the bad part" and retrain.
Solution … Every contribution gets a provenance ID
Step 1: Contributors onboard through a "Provenance-aware contribution portal"
Each dataset contribution is uploaded through a guided interface or API, is hashed and registered as an artifact, is given contributor identity (or pseudonymous handle, but stable), license terms, description and scope, optional jurisdictional data (where relevant). Each contribution artifact receives a SOVEREIGN\PROVENANCE ID and an anchor (OTS/IPFS) for external verification.
Step 2: Aggregate datasets reference contributions
When lab curates a training dataset, they build it out of contribution IDs. Sovereign composes a dataset passport: listing all constituent contribution IDs, linking to contributors and licenses, storing hash and checksum of the combined dataset.
Step 3: Training runs log dataset passports
When a training job runs, it logs the dataset passport ID(s), hyperparameters, code commit hash, environment metadata. Sovereign generates a training run record with links back to dataset passport(s) and anchors the record externally.
Flow … What happens when they suspect poisoning?
Step 1: Report
Community reports a failure mode or harmful behavior.
Step 2: Investigation
OpenLens selects the specific model version in question. Sovereign shows which training run produced it, which dataset passport that run used, which contributions are in that dataset passport.
Step 3: Inspection
They inspect contribution-level metadata: look for recent or suspicious contributions, cross-check patterns (e.g. contributions from a new, untrusted handle).
Step 4: Remediation
They can isolate a suspect contribution subset, create a new dataset passport excluding those contributions, retrain or fine-tune a corrected model.
Step 5: Transparency
They publish a public postmortem with a Sovereign provenance link showing which contributions were removed, how the retrained model's lineage differs, what controls are now in place.
"Instead of shrugging and saying 'we're not sure where that came from,' the lab can show a concrete data lineage and remediation path."
Outcomes … Trustworthy open labs instead of chaotic ones
- The lab becomes more credible to funders, more trustworthy to contributors and downstream users.
- Internally: they know they can reverse bad inputs, they can invite external review with actual structure.
- Externally: they can share provenance views as "here is how we keep open contributions responsible."
"We didn't stop being open. We made open contributions accountable."
Product tie-in … Why this is SOVEREIGN\\PROVENANCE, not just git logs
Git logs track code. MLOps tools track runs. SOVEREIGN\\PROVENANCE tracks contributions as first-class artifacts, dataset passports as composed lineage, model + run records anchored with external proofs. It ties community contributions and scientific rigor together.
Open models need open provenance. That's what Sovereign gives you.
Make every dataset contribution traceable.
If your lab runs on open data, SOVEREIGN\PROVENANCE keeps it accountable.
Why this matters now
Cases like this are emerging because AI, synthetic media, and global information flows have made provenance a necessity rather than a luxury. When it's no longer obvious who made what, when, or why, stories like this become the norm.
The Genesis Moment explains the broader context for this kind of provenance work. Read Genesis.
Work like this, when implemented on the SOVEREIGN\PROVENANCE platform, is governed by The Covenant of Restraint, our ethical framework for how we build and deploy provenance systems.
Governed by The Covenant of Restraint, aligned with our Mission & Vision, and grounded in the Genesis Moment.