The difference between open weight vs open source AI is how much of the model you actually get. An open-weight model gives you the trained parameters, so you can download, run and usually fine-tune it, but not necessarily the training data or the full training code. An open-source AI model, under the Open Source Initiative's definition, must also come with enough information about the training data and the complete code to rebuild it, all under terms that let anyone use, study, modify and share it. Most famous "open" models, including Llama, Qwen, DeepSeek, Gemma and gpt-oss, are open-weight. Fully open models such as Ai2's OLMo go further.
The licence attached to the weights matters just as much. Some open-weight models use standard permissive licences like Apache 2.0 or MIT; others use custom licences with extra conditions.
Definitions in plain English
Open weights
"Weights" are the learned numbers inside a neural network. An open-weight release publishes those numbers, plus usually the architecture and inference code, so you can run the model on your own hardware. The Open Source Initiative has an explainer on open weights noting that sharing the final parameters gives some insight into how a network operates, but reveals only a fraction of what is needed for full accountability.
Open-source AI
The Open Source AI Definition 1.0, published by the OSI in October 2024, says an open-source AI system must grant the freedoms to use it for any purpose, study how it works, modify it, and share it. To make that possible, it requires access to three things:
- Data information: a sufficiently detailed description of the training data that a skilled person could build a substantially equivalent system, including where to obtain publicly available data.
- Code: the complete source code used to train and run the system, including data processing and filtering.
- Parameters: the model weights, under OSI-approved terms.
Notice that the definition does not require publishing every byte of training data. It requires detailed data information and, where data is public or obtainable, where to get it.
The spectrum of openness
It helps to think in levels rather than two boxes.
- API only: you can call the model, but not download it.
- Open weights, restrictive licence: downloadable, with conditions on use, scale or naming.
- Open weights, permissive licence: downloadable under Apache 2.0 or MIT, with training data and code mostly private.
- Fully open: weights, training data, training code, intermediate checkpoints and logs, all released.
Licence check: popular open models in October 2026
We checked each licence on the official model card or licence file on October 5, 2026. Licences can change between versions, so always read the one attached to the exact model you download. This is a summary, not legal advice.
| Model family | Latest version checked | Weights licence | Training data released? | Openness level |
|---|---|---|---|---|
| OLMo (Ai2) | Olmo 3 and 3.1 | Apache 2.0 | Yes, plus code and checkpoints | Fully open |
| Pythia (EleutherAI) | Pythia suite | Apache 2.0 | Yes, trained on the Pile | Fully open |
| Qwen (Alibaba) | Qwen3.8-27B | Apache 2.0 | No | Open weights, permissive |
| Gemma (Google) | Gemma 4 | Apache 2.0 | No | Open weights, permissive |
| gpt-oss (OpenAI) | gpt-oss-20b and 120b | Apache 2.0, plus a usage policy | No | Open weights, permissive |
| DeepSeek | DeepSeek-V4 Flash | MIT | No | Open weights, permissive |
| Mistral | Small 4 and Large 3 | Apache 2.0 | No | Open weights, permissive |
| Mistral | Medium 3.5 | Custom licence | No | Open weights, custom terms |
| Llama (Meta) | Llama 4 | Llama 4 Community License | No | Open weights, custom terms |
A few details worth knowing:
- Gemma 4 changed the picture. Earlier Gemma models used Google's own terms. Google's open source blog says Gemma 4 is the first Gemma release under the OSI-approved Apache 2.0 licence.
- Llama 4 has notable conditions. The Llama 4 licence requires companies whose products had more than 700 million monthly active users on the release date to request a separate licence from Meta. It also requires displaying "Built with Llama" when you distribute it, and starting the name of any distributed model trained with Llama outputs with "Llama". The weights are gated on Hugging Face.
- Permissive does not mean fully open. Apache 2.0 and MIT are OSI-approved licences for the weights, but without training data information and code, these models still fall short of the OSAID's full requirements.
You can browse these families in our directory: OLMo, Pythia, Qwen, Gemma, gpt-oss, DeepSeek and Llama. Ai2 also publishes its pretraining data, including the Dolma corpus.
Why the difference matters in practice
For running models locally
For most people running a model on a laptop, any open-weight release is enough. You can quantize it to GGUF, run it in Ollama or LM Studio, and keep your data private. See our guides to running models with Ollama, llama.cpp or LM Studio, to GGUF quantization, and to running gpt-oss locally.
For building products
The licence decides what you can ship. Apache 2.0 and MIT are well understood by legal teams and allow commercial use, modification and redistribution with simple attribution requirements. Custom licences can add naming rules, attribution banners, user-count thresholds or acceptable-use policies that your company must track. Some, like gpt-oss, pair a permissive licence with a usage policy.
For fine-tuning
Open weights are what make fine-tuning an LLM locally possible. Check whether the licence places conditions on derivatives, such as Llama's naming requirement, before you publish a fine-tuned model.
For research and auditing
Only fully open models let researchers study how training data shapes behaviour, check for benchmark contamination, reproduce training runs or study learning dynamics across checkpoints. Ai2's OLMo page highlights exactly these uses, from machine unlearning to scaling studies.
Pros and cons of each approach
Open weights only
- Pros: the strongest open models are released this way; you can run, quantize and fine-tune them; many now use permissive licences.
- Cons: you cannot inspect or reproduce training; data provenance is unclear; some licences add conditions.
Fully open
- Pros: complete transparency; reproducible research; clear data provenance; ideal for teaching and auditing.
- Cons: fewer fully open models exist, and they have generally trailed the very largest open-weight releases on capability.
Which should you choose?
- You want the most capable model you can run: pick from the permissive open-weight families, and read the licence.
- You are shipping a commercial product: prefer Apache 2.0 or MIT models to minimise legal review.
- You are doing research on training, data or interpretability: start with a fully open model like OLMo or Pythia. For interpretability tooling, see TransformerLens vs nnsight.
- You care about the OSI definition for policy or procurement reasons: check for data information and training code, not just the licence name.
Who this is for
Developers choosing a base model, founders and legal teams evaluating licences, researchers who need reproducibility, and anyone confused by "open" marketing.
FAQ
What does open weight mean?
It means the trained model parameters are published for anyone to download and run, usually with the architecture and inference code. It does not necessarily include the training data or the full training code.
Is Llama open source?
Not by the OSI's definition. Llama 4 is open-weight under the Llama 4 Community License, which includes conditions such as a separate licence requirement for companies with more than 700 million monthly active users, and its training data and code are not released.
Are Qwen and DeepSeek open source?
Their weights use permissive licences, Apache 2.0 for Qwen3.8 and MIT for DeepSeek-V4 Flash, so you can use, modify and redistribute them freely. Because the training data and full training code are not published, they are open-weight rather than fully open source under the OSAID.
Which AI models are fully open source?
Ai2's OLMo family and EleutherAI's Pythia suite are well-known examples. They release weights under Apache 2.0 along with training data, training code and checkpoints.
Can I use open-weight models commercially?
Usually yes, but it depends on the licence. Apache 2.0 and MIT models allow commercial use; custom licences such as Llama's add conditions. Always read the licence for the exact model version you use.
Is Gemma open source?
Gemma 4's weights are released under Apache 2.0, an OSI-approved licence, which is a big change from earlier Gemma terms. Google does not publish the training data or full training code, so it is best described as an open-weight model with a permissive licence.