AI Research Lab · Project

Dataset Comparison

Which dataset trains a language model best?

← Back to AI Lab

Overview

We fix one architecture — the winner of our architecture comparison — and train it, under an identical budget, on several datasets of the same token count. Then we cross-evaluate every model on every dataset to see which training data produces the most generally capable model.

Objectives

  • Fair comparisonSame architecture, same token budget, same training recipe.
  • GeneralizationMeasure which data transfers best to other distributions.
  • Data quality signalUnderstand how dataset kind affects capability.
  • ReproducibilityTransparent, repeatable experiments.

Our Approach

  • Equal budgetsEach dataset streamed and capped to the same token count.
  • Cross-perplexity matrixTrain on dataset i, evaluate on dataset j, for all pairs.
  • Honest metricPerplexity isn't comparable across distributions — so we rank by mean off-diagonal (transfer) perplexity.
  • Shared harnessSame model code as the architecture study.

Focus & Tech

TinyStoriesWikiText-103OpenWebTextCross-perplexityReproducibility

Status: Harness open-sourced · runs pending architecture result

View Code on GitHub

Interested in this work?

We collaborate on data-centric research and evaluation.

Get in Touch