Versioning models and data
Know exactly which data and code produced every model, and recover or delete them when needed
A taste of a lesson
We found that a data export last month had duplicated rows. How do I know which models are affected?
This is exactly what lineage is for. Start from the faulty export: find the data snapshots created from it, ideally by hash or snapshot name. Then look up every model whose training record points to those snapshots, and every batch of predictions those models produced. If you have no such records, check model file dates against the export date and any training logs, and treat the result as uncertain. Fix the data in a new snapshot rather than editing the old one, retrain, and compare. Quick check: which three pieces of information should each model's record contain so this search takes minutes next time?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain why datasets and model files need versioning beyond code control
- Identify data versions reliably with content hashes and immutable snapshots
- Record lineage linking data, code, configuration, models and predictions
- Use registry stages and approvals to manage which model is live
- Plan for reproducibility limits, schema changes and deletion requests
Lesson plan
- 1 Why code versioning is not enough See where unversioned data and models cause real failures. Start
- 2 Hashes and immutable snapshots Identify data versions exactly and stop in place edits. Start
- 3 Lineage from data to predictions Trace any prediction back to its model, data, code and configuration. Start
- 4 Model registry ideas Manage model versions, stages and approvals with a clear record. Start
- 5 Reproducibility and schema changes Set realistic reproducibility goals and protect models from silent schema changes. Start
- 6 Retention and deletion Plan how deletion requests and retention rules apply to datasets and models. Start
Try asking
About this tutor
For practitioners who version their code but not their datasets or model files, and have already been bitten by it. You will learn why ordinary code version control struggles with large files, how content hashes and immutable snapshots identify data reliably, and how to record lineage so any model can be traced back to its training data and code. Then you look at model registry ideas, such as stages and approvals, the limits of reproducibility, schema changes, and the retention and deletion duties that come with personal data. Lessons are tool neutral and focus on the decisions and records that matter, whatever system your team uses.
Reviews
4.7
3 ratingsSample
- Ana L.Sample
The 'which data trained the live model' question at the start was uncomfortable and motivating. Good practical exercises throughout.
- Patrick O.Sample
Clear on hashing and lineage. The deletion lesson raised questions I need to take to our privacy lead, which I think is the point.
- Elif D.Sample
We had exactly the duplicated export problem. After this course we added a manifest hash and snapshot name to every training run. Simple and it already paid off.
About the teacher
MLOps without the ceremony: tracking, versioning, monitoring and responsible deployment
9 tutors 322 lessons taught Sample
I teach the habits that keep machine learning systems trustworthy after the notebook: tracking experiments, versioning data and models, testing, monitoring, handling incidents and documenting models honestly. I came to this from software operations, where I learned that most failures are boring and preventable, and then spent years helping small teams put models into production without drowning in tooling. I...
See Malik's profile and tutorsMore like this
Other tutors on the same or nearby topics.