It is tempting to download a dataset and start training immediately. But most model failures we investigate trace back to data problems that a few hours of inspection would have caught. Before you commit compute and credibility to a dataset, run through these ten checks.
1. Does it match the problem you are solving?
Write down the decision your model will support and the population it will serve. A dataset of London listings will not predict prices in Lisbon; English-only reviews will not capture Spanish or code-mixed sentiment. Mismatch between training data and real-world inputs is the single most common cause of disappointing results.
2. Is the schema documented?
Every field should have a name, type, unit and definition. Is price inclusive of tax? Is date in UTC or IST? Ambiguity here silently corrupts features.
3. How complete is it?
Profile missing values per column and per segment. Ten percent missing overall can hide 60 percent missing for one category — exactly the category your business may care about most.
4. Are there duplicates?
Near-duplicates inflate apparent dataset size and, worse, leak between training and test splits. Check for exact duplicates, then for fuzzy matches on titles, text or coordinates.
5. Is there label leakage?
Look for any feature that would not be available at prediction time. A refund_processed flag will predict churn beautifully in training and uselessly in production.
6. How reliable are the labels?
Ask how labels were produced. Human annotation should come with guidelines and an agreement score; programmatic labels should come with an estimated error rate. If neither exists, label a random sample of 200 records yourself and compare.
7. Is the distribution balanced — or honestly imbalanced?
Class imbalance is not a problem in itself, but you need to know about it to choose the right metrics. Accuracy on a 95/5 split tells you almost nothing.
8. When was the data collected?
Prices, job titles and language drift over time. Check the collection window and whether it covers seasonality such as festive sales. Plan how you will refresh the data once the model is live.
9. Can you legally use it for this purpose?
Confirm the licence permits your intended use, especially commercial use and redistribution. Check whether the data contains personal information and whether you have a lawful basis to process it under laws like India's DPDP Act.
10. Can you reproduce it?
Record the dataset version, source and any filtering you applied. If a model misbehaves six months from now, you will need to know exactly what it learned from.
A quick scoring template
| Check | Pass / Warn / Fail | Notes |
|---|---|---|
| Problem fit | ||
| Schema documentation | ||
| Completeness | ||
| Duplicates | ||
| Leakage | ||
| Label quality | ||
| Distribution | ||
| Freshness | ||
| Licence & privacy | ||
| Reproducibility |
Any Fail should block training until resolved; a handful of Warn results is normal but should be written down alongside the model.
Every dataset in the MachineLabs data store ships with a documented schema, sample rows and a clear licence, so most of this checklist is answered before you buy. Need something more specific? We can build a custom dataset to your exact requirements.