KrylovDESwitch language: Deutsch

Behind the frontier

3–4 months

That is how far openly available models recently trailed the best closed models, measured on the Epoch Capabilities Index. These are two separate estimates (about 3 months for the period to October 2025, 4 months for the first half of 2026) whose uncertainty ranges overlap. No trend can be read from them, in either direction. Epoch also notes the gap is more likely understated than overstated. For most process-automation work, a quarter behind the frontier makes no difference.

Source: Epoch AI, Epoch Capabilities Indexas of 2026-07-22

Memory for weights at 4-bit quantisation (Q4_K_M)

Parameters × 4.8944 bits ÷ 8 gives gigabytes. Despite the name, Q4_K_M uses about 4.9 bits per weight, not 4. Only the weights are shown. Context (the KV cache) and the operating system need memory too, so these bars are floors rather than requirements. The lines are ordinary system memory, not graphics memory: iPhone 17 Pro (12 GB), iPad Pro M5 (16 GB), MacBook Pro M5 Pro fully configured (64 GB) and a CPU server with 128 GB of DDR5. On the first three the model shares that memory with the operating system. Fitting is not the same as being fast: on a CPU it is memory bandwidth that limits the pace, not the space.

1 B
0.6 GB
3 B
1.8 GB
8 B
4.9 GB
14 B
8.6 GB
32 B
20 GB
70 B
43 GB
120 B
73 GB

Source: llama.cpp (bits per weight), Apple (device memory), own calculationas of 2026-07-22

Open weights does not mean free to use commercially

Colour groups related tasks; sphere size carries the parameter count of the largest released variant. These models span five orders of magnitude (22M to 1.6T parameters), so the scale is logarithmic. Twice the diameter does not mean twice the parameters. Hatched spheres may not be used commercially. Every licence was read from the official model card metadata on 22 July 2026; third-party mirrors frequently carry incorrect tags.

  • Text & code
  • Vision-language
  • Audio
  • Embeddings
Values as a table
ModelProviderTaskSizesLicenceCommercial
Gemma 4GoogleText31B · 26B-A4B · 12B · E4B · E2BApache 2.0Yes
gpt-ossOpenAIText117B-A5.1B · 21B-A3.6BApache 2.0Yes
Qwen3.6AlibabaText35B-A3B · 27BApache 2.0Yes
DeepSeek-V4-FlashDeepSeekText284B-A13BMITYes
GLM-5.2Z.aiText~753BMITYes
Kimi K2.6Moonshot AIText1T-A32BModified MITWith attribution
Llama 4 MaverickMetaText402B-A17B · 109B-A17BLlama 4 CommunityWith conditions
Mistral Small 4Mistral AIText119BApache 2.0Yes
Phi-4MicrosoftText14B · 5.6B · 3.8BMITYes
SmolLM3Hugging FaceText3BApache 2.0Yes
MiniMax-M3MiniMaxText428B-A23BMiniMax CommunityWith conditions
Nemotron 3 SuperNVIDIAText120B-A12B · 30B-A3BNVIDIA Open ModelWith attribution
Qwen3-VLAlibabaVision-language235B-A22B · 32B · 8B · 2BApache 2.0Yes
InternVL3.5OpenGVLabVision-language241B-A28B · 38B · 8B · 1BApache 2.0Yes
Molmo 2Ai2Vision-language8B · 7B · 4BApache 2.0Yes
SmolVLM2Hugging FaceVision-language2.2B · 500M · 256MApache 2.0Yes
Whisper large-v3OpenAISpeech recognition1.55BApache 2.0Yes
Whisper large-v3-turboOpenAISpeech recognition809MMITYes
Parakeet TDT v3NVIDIASpeech recognition627MCC-BY-4.0With attribution
Canary 1BNVIDIASpeech recognition1BCC-BY-NC-4.0No
Voxtral SmallMistral AISpeech recognition24B · 4.4B · 3BApache 2.0Yes
KokorohexgradSpeech synthesis82MApache 2.0Yes
ChatterboxResemble AISpeech synthesis~500MMITYes
OrpheusCanopy LabsSpeech synthesis3.8BApache 2.0Yes
F5-TTSSWividSpeech synthesis~330MCC-BY-NC-4.0No
Voxtral 4B TTSMistral AISpeech synthesis4BCC-BY-NC-4.0No
Qwen3-EmbeddingAlibabaEmbeddings8B · 4B · 0.6BApache 2.0Yes
BGE-M3BAAIEmbeddings568MMITYes
Nomic Embed v2Nomic AIEmbeddings475M MoE · 137MApache 2.0Yes
ModernBERTAnswer.AIEmbeddings396M · 150MApache 2.0Yes
all-MiniLM-L6-v2SBERTEmbeddings22MApache 2.0Yes
EmbeddingGemmaGoogleEmbeddings308MGemma TermsWith conditions
Jina Embeddings v3Jina AIEmbeddings572MCC-BY-NC-4.0No
Qwen3-CoderAlibabaCode480B-A35B · 80B-A3B · 30B-A3BApache 2.0Yes
Devstral Small 2Mistral AICode24BApache 2.0Yes
Devstral 2Mistral AICode123BModified MITWith conditions
Codestral 22BMistral AICode22BMNPLNo
DeepSeek-Coder-V2DeepSeekCode236B-A21B · 16B-A2.4BDeepSeek LicenseWith conditions

Source: Official model cards and repositoriesas of 2026-07-22

In short

A 70-billion-parameter model fits on a 64 GB laptop once quantised to 4 bits, and anything up to 120 billion fits an ordinary 128 GB server. Memory is rarely the obstacle. The question is how fast it has to be.

Book a call