To train a small LLM from scratch, pick a tiny transformer, from a few million to a few hundred million parameters, feed it a clean dataset such as TinyStories or a sample of FineWeb-Edu, and run a simple training script like Andrej Karpathy's nanochat. A model under 10 million parameters trained on TinyStories can learn to write fluent children's stories on a single consumer GPU, or slowly on a Mac. A GPT-2-class model with general knowledge is a bigger job: nanochat's reference run takes about two hours or less on a rented server with eight H100 GPUs. Training from scratch is a superb way to learn how language models work, but if you want a useful assistant for a task, fine-tuning an existing model is almost always the better choice.

Should you train from scratch at all?

GoalBest approach
Learn how LLMs really workTrain a small model from scratch
Research on architectures, optimisers or dataTrain small models from scratch and compare
A model that follows instructions for your businessFine-tune an open-weight model
A model for an unusual language or domain with lots of dataConsider continued pre-training of an existing model first

The reason is cost. Modern small models are trained on enormous amounts of text: Hugging Face's SmolLM3, a 3-billion-parameter model, was trained on 11 trillion tokens, according to its GitHub repository. You will not match that at home, and you do not need to in order to learn.

The three sizes that make sense

1. A tiny story model (laptop or single GPU)

The TinyStories paper showed that models with fewer than 10 million parameters, trained on a synthetic dataset of simple short stories, can produce fluent, consistent stories with nearly perfect grammar. The dataset is free on Hugging Face under a CDLA-Sharing licence. This is the best first project: quick to train, and you can see the model improve.

2. A GPT-1 or GPT-2 small sized model (one good GPU, or a short cloud rental)

Models from about 100 to 200 million parameters trained on a few billion tokens of filtered web text start to produce plausible general English. FineWeb-Edu publishes ready-made random samples of about 10, 100 and 350 billion tokens, so you can download the 10-billion-token sample instead of the full 1.3 trillion.

3. A GPT-2-grade chat model (an eight-GPU node for a few hours)

This is nanochat's target: a model that beats the original GPT-2's score on the DCLM CORE benchmark, then goes through fine-tuning so you can chat with it. Its README describes the cost at around 48 dollars at about 3 dollars per GPU-hour, with spot instances bringing it lower. Rental prices vary, so check your provider.

nanochat vs nanoGPT

nanoGPT was the go-to repository for years, but its README now says it is old and deprecated, and points to its successor. nanochat covers the whole pipeline in one small codebase: training a tokenizer, pre-training, fine-tuning, evaluation, and a chat interface. Its design has a single main dial, the depth of the transformer, which sets the width, learning rates and training length automatically so the model comes out compute-optimal.

Useful details from the nanochat README:

  • It runs on a single GPU by dropping torchrun; results are nearly identical but take about eight times longer.
  • On GPUs with less than 80 GB, lower the --device-batch-size setting until it fits.
  • A runcpu.sh script trains a heavily shrunk model on a CPU or Apple Silicon in a few tens of minutes, though the results are weak.
  • It is mostly plain PyTorch, so other accelerators may work, with some rough edges.

How much compute do you need?

A widely used rule of thumb from OpenAI's scaling laws paper is that training costs about 6 × parameters × tokens floating-point operations. DeepMind's Chinchilla paper found that, for a fixed budget, model size and training tokens should grow together; its 70-billion-parameter Chinchilla model was trained on 1.4 trillion tokens, about 20 tokens per parameter.

Worked example: a 124-million-parameter model trained on 2.5 billion tokens, roughly 20 tokens per parameter, needs about 6 × 124 million × 2.5 billion, or roughly 1.9 × 10 to the 18th operations. If your GPU sustains an effective 50 trillion operations per second during training, which is an assumption you should replace with your own measurement, that is about 10 to 11 hours. Halve the effective speed and the time doubles. Small models are often trained on far more than 20 tokens per parameter today, because a smaller, longer-trained model is cheaper to run afterwards.

Memory is rarely the problem at this scale. Training needs memory for the weights, the gradients, and the optimiser state, which for the common AdamW optimiser is two extra values per parameter, plus activations. A 124-million-parameter model fits easily on a 12 to 24 GB GPU with a modest batch size.

Step by step

  1. Set up. Install PyTorch with GPU support, then clone nanochat and install its dependencies with uv, as its README describes.
  2. Pick data. Start with TinyStories for a story model, or a FineWeb-Edu sample from FineWeb for general text. Load either with Hugging Face Datasets.
  3. Tokenize. Train or reuse a byte-pair-encoding tokenizer. A smaller vocabulary keeps tiny models efficient.
  4. Choose a size. For a first run, 6 to 12 layers is plenty. nanochat's author suggests a 12-layer model for quick five-minute experiments on an eight-GPU node, which means much longer on one GPU.
  5. Train and watch the loss. Log validation loss with a tool like Weights & Biases. If validation loss stops falling while training loss keeps dropping, you are overfitting or out of data.
  6. Sample often. Generate text every few hundred steps; watching gibberish turn into sentences is the fun part.
  7. Fine-tune for chat (optional). nanochat includes supervised fine-tuning and a reinforcement learning stage.
  8. Run it elsewhere (optional). Small custom architectures may not convert to GGUF, so expect to run them with the training code itself.

Common mistakes

  • Too little data for the model size. A big model on a small dataset just memorises it.
  • Judging too early. Loss curves are noisy; compare runs at the same number of tokens.
  • Changing many things at once. Change one setting per run and compare against a baseline.
  • Skipping evaluation. Keep a held-out validation set from the start.

Pros and cons of training from scratch

  • Pros: deep understanding of every stage; full control over data and architecture; a fully open model you can describe completely, as our open-weight vs open-source guide discusses.
  • Cons: small home-trained models are far weaker than released open-weight models; serious runs need rented GPUs; data preparation takes real effort.

Who this is for

Students, engineers and researchers who want to understand language models by building one, and anyone testing ideas about architectures, optimisers or data at small scale. If you are choosing a framework first, read PyTorch vs JAX.

FAQ

Can I train an LLM from scratch on my laptop?

Yes, a very small one. A model with a few million parameters trained on TinyStories is practical on a recent laptop or Mac, and nanochat includes a CPU and Apple Silicon script. Expect simple stories, not a general assistant.

How much does it cost to train a GPT-2 level model?

nanochat's README estimates about 48 dollars for its reference run, roughly two hours on an eight-H100 node at about 3 dollars per GPU-hour, and less on spot instances. Prices vary by provider.

Is nanoGPT still worth using?

Its README now marks it as deprecated and recommends nanochat instead. nanoGPT's code is still very readable for learning, but new projects should start with nanochat.

How much data do I need to train a small LLM?

The Chinchilla rule of thumb is about 20 tokens per parameter for compute-optimal training, so a 100-million-parameter model wants around 2 billion tokens. Tiny story models need far less because their language is simple.

Should I train from scratch or fine-tune?

Fine-tune if you want a capable model for a task; it is cheaper and far stronger. Train from scratch to learn, to do research, or when you need complete control over the training data.

What GPU do I need to train a small language model?

A recent NVIDIA GPU with 12 GB or more should handle models around 100 million parameters with a modest batch size. Larger runs are usually done on rented data-centre GPUs.