One Open Source Project a Day (No. 231): MiniMind — Train a 64M-Parameter LLM From Scratch in 2 Hours for About $0.43

MiniMind is a fully from-scratch, native-PyTorch LLM training project whose core claim is "train a 64-million-parameter language model from scratch in 2 hours for about ¥3." It covers the complete modern training pipeline — pretraining, SFT, LoRA, DPO, RLAIF (PPO/GRPO/CISPO), Agentic RL, and knowledge distillation — and spawns sibling projects for vision, multimodal, diffusion, and linear attention. 63.4k Stars, Apache-2.0 License.

·11 min read·AI Engineering

Introduction

"Great truths are simple — run a large language model's full journey from pretraining to reinforcement learning in the plainest possible PyTorch code."

This is the 231st article in the "One Open Source Project a Day" series. Today's project is MiniMind.

The common path into LLMs these days goes like this: install transformers, write a dozen lines to load a pretrained model, load a dataset, run inference — it works, and it works well. But what actually happens underneath those dozen lines often stays a black box. Frameworks like transformers/trl/peft separate developers from the underlying implementation. You know what you called, but not necessarily why it was built that way.

MiniMind sets out to fix exactly this "know-how without know-why" gap. It implements every core LLM training algorithm natively in PyTorch, from scratch, aimed at an almost provocative target: train a 64-million-parameter language model for about ¥3 (~$0.43 USD) in 2 hours. The smallest model here is roughly 1/2700th the size of GPT-3, and the whole pipeline runs on a single consumer GPU.

What makes it more than a toy demo: it covers the entire modern LLM training pipeline — pretraining, SFT, LoRA, DPO, RLAIF (PPO/GRPO/CISPO), Agentic RL, and knowledge distillation — and branches into sibling projects for vision (MiniMind-V), multimodal (MiniMind-O), diffusion models (MiniMind-dLM), and linear attention (MiniMind-Linear).

63.4k Stars, Apache-2.0 License, 408 commits.

What You Will Learn

  • How MiniMind implements a Qwen3-aligned Transformer architecture natively in PyTorch
  • The full training pipeline: pretrain → SFT → LoRA → DPO → RLAIF → Agentic RL → distillation
  • The "reward sparsity" problem: why rule-based RL rewards on tiny models tend to collapse to all-zero gradients
  • The "alignment tax" phenomenon in RL tuning: better at tool calls, more prone to fabrication on open-ended Q&A
  • Concrete numbers for training a usable language model on a single RTX 3090 for about ¥3 in 2 hours

Prerequisites

  • Basic understanding of Transformer architecture (attention, positional encoding, normalization layers)
  • Familiarity with standard PyTorch training loops
  • Optional: basic familiarity with RLHF/DPO/PPO-style alignment techniques

Project Background

What It Is

MiniMind's GitHub tagline is: "🧠 Train a 64M-parameter LLM from scratch in just 2h!" Its core philosophy is condensed into a single phrase — "great truths are simple." It's a fully from-scratch, open-source project for training ultra-small LLMs, aimed at letting an individual walk through the entire modern LLM training pipeline with minimal compute.

Team and Background

  • Author: jingyaogong
  • License: Apache License 2.0
  • Tech stack: Native PyTorch (every core algorithm implemented from scratch), compatible with transformers, trl, peft, llama.cpp, vllm, ollama, SGLang, Llama-Factory
  • Sibling projects: MiniMind-V (vision), MiniMind-O (multimodal/omni), MiniMind-dLM (diffusion, experimental), MiniMind-Linear (linear attention, experimental)

Project Stats

  • ⭐ GitHub Stars: 63,400+
  • 🍴 Forks: 8,200+
  • 👀 Watchers: 285
  • 📄 License: Apache-2.0
  • 📊 Commits: 408

What It Does

The Problem It Solves

The problem with the mainstream LLM learning path:
  pip install transformers
  ↓ a dozen lines: load model + load dataset + run inference
  ↑ it works, but the underlying implementation is fully walled off
  ↑ you know "what was called", not "why it was designed this way"
 
MiniMind's approach:
  Implement pretraining, SFT, LoRA, DPO, RLAIF — every core algorithm —
  natively in PyTorch
  ↓ no dependency on trl/peft wrapper layers
  ↓ run the entire pipeline for about ¥3 and 2 hours of training time
  ↑ developers can see the concrete implementation of every training stage
  ↑ upgrading from "I know how to call the API" to "I understand every
    step of the training pipeline"

Use Cases

  1. Learning LLM fundamentals

    • Want to understand how pretraining, SFT, DPO, and RLHF are actually implemented, not just how to call them
  2. Domain fine-tuning experiments

    • The project ships a medical LoRA fine-tuning example demonstrating domain adaptation on a small model
  3. Tool-use and Agentic capability research

    • Built-in Agentic RL training scripts for studying reinforcement learning in multi-turn tool-use scenarios
  4. Low-cost training pipeline validation

    • Validate training workflows, debug hyperparameters, or test new RL algorithm variants without access to large-scale compute
  5. A reference baseline for evaluation

    • Serve as a controlled baseline for studying the relationship between model scale and capability

Quick Start

Install and run inference:

git clone --depth 1 https://github.com/jingyaogong/minimind
cd minimind && pip install -r requirements.txt
 
# Download a model from ModelScope or HuggingFace, then run inference
python eval_llm.py --load_from ./minimind-3

Train from scratch:

# Download dataset files into ./dataset first
 
cd trainer
python train_pretrain.py      # Pretraining
python train_full_sft.py      # Full supervised fine-tuning
 
# Multi-GPU training
torchrun --nproc_per_node N train_xxx.py

Deployment:

# Run directly with ollama
ollama run jingyaogong/minimind-3
 
# Or deploy with vllm
vllm serve ...
 
# Or launch the Streamlit WebUI / OpenAI-compatible API server

Core Features

1. Model Lineup

ModelParamsLayersd_modelNotes
minimind-364M8768Dense
minimind-3-moe198M-A64M87684 experts, Top-1 routing
minimind2-small26M8512Historical
minimind2104M16768Historical

2. Architecture Design

Transformer Decoder-Only, aligned with the Qwen3 ecosystem for compatibility with transformers/llama.cpp/ollama/vllm:

  • Pre-Normalization + RMSNorm
  • SwiGLU activation function
  • RoPE positional encoding with YaRN long-context extrapolation support
  • Attention config: q_heads=8, kv_heads=4, max_position_embeddings=32768, rope_theta=1e6

3. A Deliberately Compact Vocabulary

The tokenizer vocabulary is just 6,400 tokens (compared to Qwen2's 151,643 or Llama 3's 128,000). This is a deliberate tradeoff — at this parameter scale, the embedding and output layers' share of total parameters has to be squeezed down as much as possible.

4. Full Training Pipeline

StagePurpose
PretrainingNext-token prediction on large-scale text corpora (unsupervised)
SFTSupervised fine-tuning on multi-turn dialogue, tool calls, and thinking-tag formatting
LoRAParameter-efficient fine-tuning implemented from scratch (no peft dependency)
DPOPreference optimization via chosen/rejected pairs
RLAIFPPO, GRPO, CISPO — reinforcement learning with AI/rule-based feedback
Agentic RLMulti-turn tool-use training via train_agent.py
Knowledge DistillationBlack-box distillation (teacher-output fine-tuning) + white-box distillation (CE+KL distribution matching)

5. Measured Training Cost

ConfigurationTimeCost
Full pipeline (pretrain_mini + sft_mini), single RTX 3090~2.31 hours¥3.0 ($0.43)
MoE variant~3.23 hours¥4.2 ($0.60)
8×H100Reduced to "minutes"—

A Deeper Look

Unifying the RL Algorithm Family Under One Lens

One design choice in MiniMind worth examining is how it frames its RL training stage. Rather than treating PPO, GRPO, and CISPO as three isolated algorithms implemented separately, it frames them under a single unified view: all policy optimization algorithms are natural variants formed by making different design trade-offs on the same objective function — specifically on the policy term, the advantage term, and the regularization term.

The unified view of PO algorithms:
  Policy term
  Advantage term
  Regularization term
  ↓
  PPO, GRPO, and CISPO are essentially different concrete implementation
  choices on these three terms
  ↑ not three isolated algorithms, but parameterized variants of the
    same framework

The practical value of this framing: it collapses half the memorization burden. Once you understand the design space of these three terms, you understand the entire RL alignment algorithm family — instead of memorizing each algorithm as an isolated black box.

"Reward Sparsity": The Hidden Trap of RL on Tiny Models

MiniMind's documentation surfaces a very practical engineering lesson — using rule-based/binary rewards for RL training on small models tends to produce a "reward sparsity" problem on hard tasks. The project uses a vivid analogy: this is like making an elementary school student take college entrance exam math problems.

The problem scenario:
  A small model almost never gets the correct answer on hard tasks
  ↓ rule-based reward (correct = 1, incorrect = 0)
  ↓ nearly every rollout scores 0
  ↑ the gradient signal collapses to zero, and the policy can't learn
 
MiniMind's mitigation:
  Use continuous reward-model scores (InternLM2-1.8B-Reward) instead of
  binary rule-based rewards
  ↑ even a partially correct answer provides a gradual gradient signal
  ↑ mitigates reward sparsity on hard tasks for tiny models

This issue tends to be overlooked in large-model RL training, because large models already have a high enough task success rate that rule-based rewards provide sufficient gradient signal on their own. But for an extremely small model like 64M parameters, reward sparsity is a problem that has to be explicitly engineered around — and it's a hands-on lesson rarely spelled out in large-model project documentation.

Measured Gains From Agentic RL, and the "Alignment Tax"

The project provides a concrete tool-use comparison: on 20 math tool-call tasks, a model trained with SFT only scored 12/20 (60%), while the same model after Agent RL training scored 17/20 (85%). This is a clean, quantified piece of evidence that Agentic RL training genuinely improves tool-use capability.

But the project is equally candid about the cost — RL-tuned models exhibit an "alignment tax": they perform better on the specific task they were tuned for (tool calling) but become "more prone to fabrication" on open-ended Q&A. In other words, RL training makes the model more reliable within its training distribution, but potentially less reliable outside of it. This kind of candid limitation disclosure is relatively rare in project documentation that usually emphasizes performance gains.

Comparison With Similar Small Chinese LLM Projects

The project directly compares itself to baby-llama2-chinese (0.2B) and chatlm-mini-chinese (0.2B), but chooses qualitative output comparison rather than standardized leaderboard scores — showing real Q&A examples and observing whether MiniMind's output is fluent, and whether it shows factual inconsistency or repetition.

Dimensionbaby-llama2-chinese / chatlm-mini-chineseMiniMind
Training pipeline coverageUsually stops at pretrain + SFTFull pretrain→SFT→LoRA→DPO→RLAIF→Agentic RL→distillation
Architecture alignmentIndependently designedAligned with Qwen3/Qwen3-MoE ecosystem, compatible with mainstream inference frameworks
Agentic capabilityUsually not coveredNative train_agent.py, supports GRPO/CISPO multi-turn tool-use training
Training cost transparencyRarely disclosed with specific numbersExplicit RTX 3090 time and RMB cost figures
Limitation disclosureRareProactively documents side effects like "alignment tax"

MiniMind's differentiation isn't "smaller parameter count" or "higher benchmark scores" — it's that it shows you, with the plainest possible implementation, every stage of the modern LLM training pipeline, including the unglamorous engineering lessons (reward sparsity, alignment tax) that most projects leave out.


Official Resources

  • MiniMind-V — MiniMind's vision/multimodal sibling project
  • Qwen3 — the ecosystem MiniMind's architecture is aligned with
  • InternLM2-1.8B-Reward — the default reward model used in the RLAIF stage

Summary

Key Takeaways

  1. Native PyTorch implementation pries open the training black box: no dependency on trl/peft wrapper layers; every core algorithm built from scratch
  2. Low-cost and reproducible: trains a usable 64M-parameter model on a single RTX 3090 for about ¥3 in 2 hours
  3. Covers the full modern LLM training pipeline: pretrain→SFT→LoRA→DPO→RLAIF→Agentic RL→distillation, end to end
  4. A unified lens for understanding the RL algorithm family: PPO/GRPO/CISPO as different trade-offs on policy/advantage/regularization terms within the same objective
  5. Candidly discloses its own limitations: openly documents real engineering problems like reward sparsity and alignment tax, not just the highlights

Who This Is For

  • Learners who want to deeply understand LLM training mechanics: not satisfied with calling an API, want to see every step from pretraining to RL alignment implemented concretely
  • Individual developers/researchers with limited compute: want to validate training pipelines or test new algorithm variants on a consumer GPU
  • Anyone researching domain fine-tuning or Agentic applications: can directly reuse its LoRA, DPO, and Agentic RL implementation patterns
  • Instructors/students of large-model courses: MiniMind's code size and complexity make it well-suited as a teaching example

One-Line Verdict

MiniMind proves that understanding how a large language model is actually trained doesn't require dozens of A100s — just one RTX 3090, about $0.43, and the patience to read every line of training code.


Check out PrimeSkills — a curated marketplace of AI agents and skills that have been validated in real-world, enterprise-grade workflows. No fluff, just what actually works.

Find more useful knowledge and interesting products on my Homepage