Notes
Open models in practice.
Three questions decide whether a model can run in your own building: how far it trails the frontier, how much memory it needs, and whether you are allowed to use it commercially at all.
Behind the frontier
3–4 months
That is how far openly available models recently trailed the best closed models, measured on the Epoch Capabilities Index. These are two separate estimates (about 3 months for the period to October 2025, 4 months for the first half of 2026) whose uncertainty ranges overlap. No trend can be read from them, in either direction. Epoch also notes the gap is more likely understated than overstated. For most process-automation work, a quarter behind the frontier makes no difference.
Source: Epoch AI, Epoch Capabilities Indexas of 2026-07-22
Memory for weights at 4-bit quantisation (Q4_K_M)
Parameters × 4.8944 bits ÷ 8 gives gigabytes. Despite the name, Q4_K_M uses about 4.9 bits per weight, not 4. Only the weights are shown. Context (the KV cache) and the operating system need memory too, so these bars are floors rather than requirements. The lines are ordinary system memory, not graphics memory: iPhone 17 Pro (12 GB), iPad Pro M5 (16 GB), MacBook Pro M5 Pro fully configured (64 GB) and a CPU server with 128 GB of DDR5. On the first three the model shares that memory with the operating system. Fitting is not the same as being fast: on a CPU it is memory bandwidth that limits the pace, not the space.
Source: llama.cpp (bits per weight), Apple (device memory), own calculationas of 2026-07-22
Open weights does not mean free to use commercially
Colour groups related tasks; sphere size carries the parameter count of the largest released variant. These models span five orders of magnitude (22M to 1.6T parameters), so the scale is logarithmic. Twice the diameter does not mean twice the parameters. Hatched spheres may not be used commercially. Every licence was read from the official model card metadata on 22 July 2026; third-party mirrors frequently carry incorrect tags.
- Text & code
- Vision-language
- Audio
- Embeddings
Kimi K2.6
- Provider
- Moonshot AI
- Sizes
- 1T-A32B
- Licence
- Modified MIT
- Commercial
- With attribution
Attribution only kicks in above 100M users or $20M monthly revenue.
GLM-5.2
- Provider
- Z.ai
- Sizes
- ~753B
- Licence
- MIT
- Commercial
- Yes
MiniMax-M3
- Provider
- MiniMax
- Sizes
- 428B-A23B
- Licence
- MiniMax Community
- Commercial
- With conditions
A one-time email notice is required; above $20M annual revenue you need written authorisation.
Llama 4 Maverick
- Provider
- Meta
- Sizes
- 402B-A17B · 109B-A17B
- Licence
- Llama 4 Community
- Commercial
- With conditions
Above 700M monthly active users a separate licence is required. "Built with Llama" must be displayed.
DeepSeek-V4-Flash
- Provider
- DeepSeek
- Sizes
- 284B-A13B
- Licence
- MIT
- Commercial
- Yes
Nemotron 3 Super
- Provider
- NVIDIA
- Sizes
- 120B-A12B · 30B-A3B
- Licence
- NVIDIA Open Model
- Commercial
- With attribution
Mistral Small 4
- Provider
- Mistral AI
- Sizes
- 119B
- Licence
- Apache 2.0
- Commercial
- Yes
gpt-oss
- Provider
- OpenAI
- Sizes
- 117B-A5.1B · 21B-A3.6B
- Licence
- Apache 2.0
- Commercial
- Yes
Qwen3.6
- Provider
- Alibaba
- Sizes
- 35B-A3B · 27B
- Licence
- Apache 2.0
- Commercial
- Yes
Gemma 4
- Provider
- Sizes
- 31B · 26B-A4B · 12B · E4B · E2B
- Licence
- Apache 2.0
- Commercial
- Yes
New with Gemma 4. Gemma 3 and earlier remain under the Gemma Terms.
Phi-4
- Provider
- Microsoft
- Sizes
- 14B · 5.6B · 3.8B
- Licence
- MIT
- Commercial
- Yes
SmolLM3
- Provider
- Hugging Face
- Sizes
- 3B
- Licence
- Apache 2.0
- Commercial
- Yes
Qwen3-Coder
- Provider
- Alibaba
- Sizes
- 480B-A35B · 80B-A3B · 30B-A3B
- Licence
- Apache 2.0
- Commercial
- Yes
DeepSeek-Coder-V2
- Provider
- DeepSeek
- Sizes
- 236B-A21B · 16B-A2.4B
- Licence
- DeepSeek License
- Commercial
- With conditions
Carries use restrictions, unlike DeepSeek V3, R1 and V4, which are MIT.
Devstral 2
- Provider
- Mistral AI
- Sizes
- 123B
- Licence
- Modified MIT
- Commercial
- With conditions
Not licensed above $20M monthly revenue. The smaller Devstral Small 2 is plain Apache 2.0.
Devstral Small 2
- Provider
- Mistral AI
- Sizes
- 24B
- Licence
- Apache 2.0
- Commercial
- Yes
Codestral 22B
- Provider
- Mistral AI
- Sizes
- 22B
- Licence
- MNPL
- Commercial
- No
Mistral Non-Production Licence: no production use, despite Apache 2.0 siblings in the same org.
InternVL3.5
- Provider
- OpenGVLab
- Sizes
- 241B-A28B · 38B · 8B · 1B
- Licence
- Apache 2.0
- Commercial
- Yes
Qwen3-VL
- Provider
- Alibaba
- Sizes
- 235B-A22B · 32B · 8B · 2B
- Licence
- Apache 2.0
- Commercial
- Yes
Molmo 2
- Provider
- Ai2
- Sizes
- 8B · 7B · 4B
- Licence
- Apache 2.0
- Commercial
- Yes
SmolVLM2
- Provider
- Hugging Face
- Sizes
- 2.2B · 500M · 256M
- Licence
- Apache 2.0
- Commercial
- Yes
Voxtral Small
- Provider
- Mistral AI
- Sizes
- 24B · 4.4B · 3B
- Licence
- Apache 2.0
- Commercial
- Yes
Whisper large-v3
- Provider
- OpenAI
- Sizes
- 1.55B
- Licence
- Apache 2.0
- Commercial
- Yes
Canary 1B
- Provider
- NVIDIA
- Sizes
- 1B
- Licence
- CC-BY-NC-4.0
- Commercial
- No
The original release only. Canary 1B v2 and Canary-Qwen are CC-BY-4.0.
Whisper large-v3-turbo
- Provider
- OpenAI
- Sizes
- 809M
- Licence
- MIT
- Commercial
- Yes
Parakeet TDT v3
- Provider
- NVIDIA
- Sizes
- 627M
- Licence
- CC-BY-4.0
- Commercial
- With attribution
Voxtral 4B TTS
- Provider
- Mistral AI
- Sizes
- 4B
- Licence
- CC-BY-NC-4.0
- Commercial
- No
The only Voxtral model that is not Apache 2.0: same family name, opposite answer.
Orpheus
- Provider
- Canopy Labs
- Sizes
- 3.8B
- Licence
- Apache 2.0
- Commercial
- Yes
Chatterbox
- Provider
- Resemble AI
- Sizes
- ~500M
- Licence
- MIT
- Commercial
- Yes
F5-TTS
- Provider
- SWivid
- Sizes
- ~330M
- Licence
- CC-BY-NC-4.0
- Commercial
- No
Very widely used and still not commercially usable. Derived fine-tunes inherit the restriction.
Kokoro
- Provider
- hexgrad
- Sizes
- 82M
- Licence
- Apache 2.0
- Commercial
- Yes
Qwen3-Embedding
- Provider
- Alibaba
- Sizes
- 8B · 4B · 0.6B
- Licence
- Apache 2.0
- Commercial
- Yes
Jina Embeddings v3
- Provider
- Jina AI
- Sizes
- 572M
- Licence
- CC-BY-NC-4.0
- Commercial
- No
The entire current Jina line (v3, v5, rerankers) is non-commercial. Only the older v2 models are Apache 2.0.
BGE-M3
- Provider
- BAAI
- Sizes
- 568M
- Licence
- MIT
- Commercial
- Yes
Nomic Embed v2
- Provider
- Nomic AI
- Sizes
- 475M MoE · 137M
- Licence
- Apache 2.0
- Commercial
- Yes
ModernBERT
- Provider
- Answer.AI
- Sizes
- 396M · 150M
- Licence
- Apache 2.0
- Commercial
- Yes
EmbeddingGemma
- Provider
- Sizes
- 308M
- Licence
- Gemma Terms
- Commercial
- With conditions
Did not move to Apache 2.0 the way Gemma 4 did. Licence acceptance required.
all-MiniLM-L6-v2
- Provider
- SBERT
- Sizes
- 22M
- Licence
- Apache 2.0
- Commercial
- Yes
The most-downloaded model on Hugging Face, full stop.
Values as a table
| Model | Provider | Task | Sizes | Licence | Commercial |
|---|---|---|---|---|---|
| Gemma 4 | Text | 31B · 26B-A4B · 12B · E4B · E2B | Apache 2.0 | Yes | |
| gpt-oss | OpenAI | Text | 117B-A5.1B · 21B-A3.6B | Apache 2.0 | Yes |
| Qwen3.6 | Alibaba | Text | 35B-A3B · 27B | Apache 2.0 | Yes |
| DeepSeek-V4-Flash | DeepSeek | Text | 284B-A13B | MIT | Yes |
| GLM-5.2 | Z.ai | Text | ~753B | MIT | Yes |
| Kimi K2.6 | Moonshot AI | Text | 1T-A32B | Modified MIT | With attribution |
| Llama 4 Maverick | Meta | Text | 402B-A17B · 109B-A17B | Llama 4 Community | With conditions |
| Mistral Small 4 | Mistral AI | Text | 119B | Apache 2.0 | Yes |
| Phi-4 | Microsoft | Text | 14B · 5.6B · 3.8B | MIT | Yes |
| SmolLM3 | Hugging Face | Text | 3B | Apache 2.0 | Yes |
| MiniMax-M3 | MiniMax | Text | 428B-A23B | MiniMax Community | With conditions |
| Nemotron 3 Super | NVIDIA | Text | 120B-A12B · 30B-A3B | NVIDIA Open Model | With attribution |
| Qwen3-VL | Alibaba | Vision-language | 235B-A22B · 32B · 8B · 2B | Apache 2.0 | Yes |
| InternVL3.5 | OpenGVLab | Vision-language | 241B-A28B · 38B · 8B · 1B | Apache 2.0 | Yes |
| Molmo 2 | Ai2 | Vision-language | 8B · 7B · 4B | Apache 2.0 | Yes |
| SmolVLM2 | Hugging Face | Vision-language | 2.2B · 500M · 256M | Apache 2.0 | Yes |
| Whisper large-v3 | OpenAI | Speech recognition | 1.55B | Apache 2.0 | Yes |
| Whisper large-v3-turbo | OpenAI | Speech recognition | 809M | MIT | Yes |
| Parakeet TDT v3 | NVIDIA | Speech recognition | 627M | CC-BY-4.0 | With attribution |
| Canary 1B | NVIDIA | Speech recognition | 1B | CC-BY-NC-4.0 | No |
| Voxtral Small | Mistral AI | Speech recognition | 24B · 4.4B · 3B | Apache 2.0 | Yes |
| Kokoro | hexgrad | Speech synthesis | 82M | Apache 2.0 | Yes |
| Chatterbox | Resemble AI | Speech synthesis | ~500M | MIT | Yes |
| Orpheus | Canopy Labs | Speech synthesis | 3.8B | Apache 2.0 | Yes |
| F5-TTS | SWivid | Speech synthesis | ~330M | CC-BY-NC-4.0 | No |
| Voxtral 4B TTS | Mistral AI | Speech synthesis | 4B | CC-BY-NC-4.0 | No |
| Qwen3-Embedding | Alibaba | Embeddings | 8B · 4B · 0.6B | Apache 2.0 | Yes |
| BGE-M3 | BAAI | Embeddings | 568M | MIT | Yes |
| Nomic Embed v2 | Nomic AI | Embeddings | 475M MoE · 137M | Apache 2.0 | Yes |
| ModernBERT | Answer.AI | Embeddings | 396M · 150M | Apache 2.0 | Yes |
| all-MiniLM-L6-v2 | SBERT | Embeddings | 22M | Apache 2.0 | Yes |
| EmbeddingGemma | Embeddings | 308M | Gemma Terms | With conditions | |
| Jina Embeddings v3 | Jina AI | Embeddings | 572M | CC-BY-NC-4.0 | No |
| Qwen3-Coder | Alibaba | Code | 480B-A35B · 80B-A3B · 30B-A3B | Apache 2.0 | Yes |
| Devstral Small 2 | Mistral AI | Code | 24B | Apache 2.0 | Yes |
| Devstral 2 | Mistral AI | Code | 123B | Modified MIT | With conditions |
| Codestral 22B | Mistral AI | Code | 22B | MNPL | No |
| DeepSeek-Coder-V2 | DeepSeek | Code | 236B-A21B · 16B-A2.4B | DeepSeek License | With conditions |
Source: Official model cards and repositoriesas of 2026-07-22
In short
A 70-billion-parameter model fits on a 64 GB laptop once quantised to 4 bits, and anything up to 120 billion fits an ordinary 128 GB server. Memory is rarely the obstacle. The question is how fast it has to be.
Book a call