After downloading the model, you may encounter an out-of-memory or out-of-video-memory error at runtime. This is probably the last thing anyone wants to see when trying a local model.
As introduced earlier, the model needs to load its weights into RAM or VRAM at runtime, and generating responses also requires additional memory. If your device cannot accommodate the model, aside from switching to a smaller model, you can also check whether a quantized version is available.
Model weights consist of a large number of numerical values, and storing each value requires a certain amount of space. Quantization reduces the model's resource consumption by lowering the precision of these numerical representations. For example, weights originally stored as 16-bit floating-point numbers can be converted to predominantly 8-bit or 4-bit low-precision representations. The number of parameters usually remains the same, but the space occupied by each parameter becomes smaller.
Taking a 7 billion parameter model as an example, and only considering the weights themselves without quantization overhead, we can get the following rough estimates:
Weight Representation | Theoretical Weight Size |
16 bit | Approximately 14 GB |
8 bit | Approximately 7 GB |
4 bit | Approximately 3.5 GB |
The GB values here are calculated in decimal units. Actual quantized files also store information such as scaling factors and may use mixed precision, so the actual size may differ from the estimate.
For ordinary users, the most direct benefit of quantization is that model files become smaller, saving space for downloading and storage, and requiring less RAM or VRAM for weights at runtime. When supported by hardware and inference frameworks, quantization may also improve inference speed.
However, reducing the file to one-quarter of its original size does not mean the response speed will increase fourfold.
Lowering precision may also affect response quality, depending on the model, quantization method, and task. When choosing a quantized version, besides confirming it can run, you should also test it with your own questions to see if the responses remain reliable.
For models that have already been downloaded or fine-tuned, the common approach is Post-Training Quantization (PTQ), which converts weights and other numerical values to lower precision representations after model training is complete. Some methods require a small batch of representative data to calibrate the quantization process and minimize error.
Another category is , which simulates the effects of quantization during training, allowing the model to adapt to low precision representations in advance.
Sign in to join the discussion
When running models with Ollama, you will frequently encounter names like GGUF, Q4, Q8, etc. These names are mainly related to the model's storage format and quantization method, and are a key reason why large models can run on personal computers.
In August 2023, Georgi Gerganov introduced the GGUF format. Its main advantages include single-file deployment, extensibility, support for memory mapping, complete information, and ease of use. A GGUF file includes a file header, metadata key-value pairs, and tensor information. These components together define the model's structure and behavior. As shown below, GGUF is a model file format for large model inference, currently widely used by local inference tools such as llama.cpp and Ollama. Ollama encapsulates GGUF format handling, model quantization, and model loading processes, so users don't need to convert model formats themselves โ they can directly download and run pre-quantized models. A GGUF file contains more than just model parameters; it also includes basic model information, model architecture, and tensor data. GGUF organizes this information through a unified file format, allowing inference tools like llama.cpp to directly read and load it. Its basic structure is shown in the figure below.
ModelScope also supports displaying structured GGUF files, as shown below. When inference tools load GGUF files, they can directly read this information and load the model without needing to prepare multiple model configuration files separately.
GGUF itself is just a model file format and does not directly reduce the model's RAM or VRAM usage. What truly reduces model resource consumption is quantization.
In llama.cpp and GGUF models, Q4, Q5, Q8, Q4_K_M represent different quantization names. Here, Q stands for quantization, and the following number indicates the primary quantization bit width. For example, Q4 means primarily 4-bit quantization, and Q8 means primarily 8-bit quantization.
Note that Q4 does not mean all parameters in the model uniformly use 4-bit representation. Taking the common Q4_K_M as an example, it belongs to the K-quant quantization type in llama.cpp. The Q4 in the name means primarily 4-bit quantization, K indicates the use of the K-quant quantization scheme, and M denotes the Medium configuration within that scheme. In practice, different tensors in the model may use different quantization precisions rather than all parameters uniformly using 4-bit representation, as shown in the figure below.
When choosing a GGUF model, first confirm that your inference tool supports the model, then look at the specific quantization version. Lower quantization precision generally reduces RAM and VRAM usage, making it easier to run the model on ordinary devices, but may sacrifice some model performance. Higher quantization precision generally preserves more of the original model's capabilities but also requires more hardware resources.
If a version is already close to your device's RAM or VRAM limits, consider switching to a smaller version to leave room for context and runtime operations. If resources are sufficient, you can use the same questions to compare different versions' responses, paying special attention to your target tasks such as information extraction, code generation, or structured output.
First confirm stable operation, then check response quality. This way, the quantized version you choose will be better suited to your device and use case.