AI Research Lab · Project
Dataset Comparison
Which dataset trains a language model best?
Overview
We fix one architecture — the winner of our architecture comparison — and train it, under an identical budget, on several datasets of the same token count. Then we cross-evaluate every model on every dataset to see which training data produces the most generally capable model.
Objectives
- Fair comparisonSame architecture, same token budget, same training recipe.
- GeneralizationMeasure which data transfers best to other distributions.
- Data quality signalUnderstand how dataset kind affects capability.
- ReproducibilityTransparent, repeatable experiments.
Our Approach
- Equal budgetsEach dataset streamed and capped to the same token count.
- Cross-perplexity matrixTrain on dataset i, evaluate on dataset j, for all pairs.
- Honest metricPerplexity isn't comparable across distributions — so we rank by mean off-diagonal (transfer) perplexity.
- Shared harnessSame model code as the architecture study.
Focus & Tech
TinyStoriesWikiText-103OpenWebTextCross-perplexityReproducibility
Status: Harness open-sourced · runs pending architecture result
Interested in this work?
We collaborate on data-centric research and evaluation.