// guides · 15 articles
Guides
Original, hands-on guides from the Axiomi editors: comparisons, “best tool for the job” round-ups and plain-English explainers.
Comparisons · 7 min readListen
vLLM vs llama.cpp: Which Inference Engine Fits?
vLLM vs llama.cpp compared, with SGLang and Ollama: throughput vs single-user speed, hardware, model formats, quantization, ports and when to pick each engine.
Comparisons · 6 min readListen
TransformerLens vs nnsight: Which Interp Library?
TransformerLens vs nnsight for mechanistic interpretability: the TransformerLens 4.0 bridge, nnsight tracing, free NDIF remote runs, code samples, how to pick.
Guides · 6 min readListen
Train a Small LLM From Scratch: A Practical Guide
How to train a small LLM from scratch in 2026: realistic sizes and budgets, nanochat vs nanoGPT, FineWeb-Edu and TinyStories data, the compute math and steps.
Explainers · 6 min readListen
Speculative Decoding vs MTP: Faster LLM Output
Speculative decoding vs MTP explained: how draft models, EAGLE-3, MTP heads and DFlash speed up LLM output, and how to enable each in llama.cpp, vLLM or SGLang.
Comparisons · 6 min readListen
PyTorch vs JAX in 2026: Which Should You Learn?
PyTorch vs JAX in 2026: how eager PyTorch and functional JAX differ, hardware support, the LLM ecosystem, TPUs and TorchTPU, and which framework fits your work.
Explainers · 6 min readListen
Open Weight vs Open Source AI: The Difference
Open weight vs open source AI explained: what the OSI definition requires, which popular models are which, and a licence table checked against model cards.
Comparisons · 8 min readListen
Ollama vs llama.cpp vs LM Studio: 2026 Guide
Ollama vs llama.cpp vs LM Studio compared for 2026: what each one is, ports, licences, Mac and GPU support, the context-length trap, and which to pick.
Comparisons · 7 min readListen
MLX vs GGUF on Mac: Which Should You Use?
MLX vs GGUF on a Mac: what each format is, how Ollama and LM Studio use them, speed trade-offs, memory limits, fine-tuning with mlx-lm, and which to download.
Explainers · 7 min readListen
LoRA vs QLoRA vs DoRA: Which Fine-Tuning Method?
QLoRA vs LoRA explained, plus DoRA and full fine-tuning: how each works, real VRAM figures, quality trade-offs, recommended settings, and when to choose which.
Explainers · 6 min readListen
KV Cache Quantization: Fit Longer Context in VRAM
KV cache quantization explained: how q8_0, q4_0 and FP8 caches cut memory, the settings for llama.cpp, Ollama, LM Studio, vLLM and MLX, and when quality drops.
How-to · 6 min readListen
How to Run gpt-oss Locally (20b and 120b)
Run gpt-oss locally with Ollama, LM Studio, llama.cpp or vLLM: memory needs for 20b and 120b, setup commands, reasoning effort, harmony format and fine-tuning.
How-to · 7 min readListen
How to Fine-Tune an LLM Locally in 2026
Fine-tune an LLM locally, step by step: when it pays off, picking a base model, formatting data, Unsloth vs Axolotl vs TRL vs mlx-lm, training and GGUF export.
Explainers · 7 min readListen
How Much VRAM Do I Need for an LLM?
How much VRAM do you need for an LLM? A simple formula for weights plus KV cache, worked examples from real model configs, and what fits in 8, 16, 24 or 32 GB.
Explainers · 6 min readListen
GRPO vs PPO: How LLM RL Training Differs
GRPO vs PPO explained simply: how each trains LLMs with rewards, why GRPO drops the critic, what DAPO and Dr. GRPO changed, and when to pick GRPO, PPO or DPO.
Explainers · 7 min readListen
GGUF Quantization Explained: Pick the Right Quant
GGUF quantization explained in plain English: what Q4_K_M, Q8_0 and IQ quants mean, real bits-per-weight numbers, GGUF vs safetensors, and which file to pick.