AIUnlimited
๐ŸŒณ

AI Foundations

๐ŸŒฑ
AI Seeds

Start from zero

๐ŸŒฟ
AI Sprouts

Build foundations

๐ŸŒณ
AI Branches

Apply in practice

๐Ÿ•๏ธ
AI Canopy

Go deep

๐ŸŒฒ
AI Forest

Master AI

๐Ÿ”จ

AI Mastery

โœ๏ธ
AI Sketch

Start from zero

๐Ÿชจ
AI Chisel

Build foundations

โš’๏ธ
AI Craft

Apply in practice

๐Ÿ’Ž
AI Polish

Go deep

๐Ÿ†
AI Masterpiece

Master AI

๐Ÿ“˜

AI Practice

๐Ÿ“–
Understanding Open-Source Models

Fundamentals and resources for open-source models

๐ŸŽฏ
From Problem to Model Task

Converting business problems to model tasks

โšก
Running Your First Model

See your first results in 30 minutes

๐Ÿ”ง
Fine-Tuning and Evaluation

Fine-tune models and evaluate performance

๐Ÿš€
Application Systems

Build real-world AI applications

๐ŸŽจ
Generative AI

Explore open-source AIGC models

๐Ÿค–
Agents

Learn Agent frameworks and MCP tools

๐Ÿ“
Supplementary Fundamentals

LLM basics and evaluation

๐ŸŽ“

Claude Academy

๐Ÿค–
Claude 101

Learn AI basics with Claude

๐Ÿ’ป
Claude Code 101

Code with Claude as your pair programmer

๐Ÿค
Introduction to Claude Cowork

Collaborate with Claude on complex projects

โš™๏ธ
Claude Platform 101

Build apps with the Claude API

Lab

7 experiments loaded
๐ŸงฌNeural Network Sandbox๐Ÿค–AI or Human?๐Ÿฅ‹Prompt Engineering Dojo๐ŸAlgorithm Race๐Ÿง AI Trivia Challenge๐Ÿ—๏ธSystem Design Canvas
๐ŸŽฏMock InterviewEnter the Labโ†’
๐Ÿš€

Career Development

๐Ÿš€
Interview Launchpad

Start your journey

๐ŸŒŸ
Behavioral Mastery

Master soft skills

๐Ÿ’ป
Technical Interviews

Ace the coding round

๐Ÿค–
AI & ML Interviews

ML interview mastery

๐Ÿ†
Offer & Beyond

Land the best offer

Get Started
AIUnlimited

MIT Licence.

ๆฒชICPๅค‡18025655ๅท-11

Learn

  • AI Basics
  • AI Practice
  • Claude Academy
  • Lab
  • Career Development

Community

  • About
  • FAQ

Support

  • Terms of Service
  • Privacy Policy
  • Contact
AI & Engineering Academicsโ€บโšก Running Your First Modelโ€บLessonsโ€บModel Quantization Basics
๐Ÿ“ฆ
Running Your First Model โ€ข Beginnerโฑ๏ธ 15 min read

Model Quantization Basics

Models Require Too Many Resources โ€” What Can Quantization Do?

After downloading the model, you may encounter an out-of-memory or out-of-video-memory error at runtime. This is probably the last thing anyone wants to see when trying a local model.

As introduced earlier, the model needs to load its weights into RAM or VRAM at runtime, and generating responses also requires additional memory. If your device cannot accommodate the model, aside from switching to a smaller model, you can also check whether a quantized version is available.

How Much Space Can Quantization Save?

Model weights consist of a large number of numerical values, and storing each value requires a certain amount of space. Quantization reduces the model's resource consumption by lowering the precision of these numerical representations. For example, weights originally stored as 16-bit floating-point numbers can be converted to predominantly 8-bit or 4-bit low-precision representations. The number of parameters usually remains the same, but the space occupied by each parameter becomes smaller.

Taking a 7 billion parameter model as an example, and only considering the weights themselves without quantization overhead, we can get the following rough estimates:

Weight Representation

Theoretical Weight Size

16 bit

Approximately 14 GB

8 bit

Approximately 7 GB

4 bit

Approximately 3.5 GB

The GB values here are calculated in decimal units. Actual quantized files also store information such as scaling factors and may use mixed precision, so the actual size may differ from the estimate.

For ordinary users, the most direct benefit of quantization is that model files become smaller, saving space for downloading and storage, and requiring less RAM or VRAM for weights at runtime. When supported by hardware and inference frameworks, quantization may also improve inference speed.

However, reducing the file to one-quarter of its original size does not mean the response speed will increase fourfold.

Lowering precision may also affect response quality, depending on the model, quantization method, and task. When choosing a quantized version, besides confirming it can run, you should also test it with your own questions to see if the responses remain reliable.

What Approaches Are Available for Model Quantization?

For models that have already been downloaded or fine-tuned, the common approach is Post-Training Quantization (PTQ), which converts weights and other numerical values to lower precision representations after model training is complete. Some methods require a small batch of representative data to calibrate the quantization process and minimize error.

Another category is , which simulates the effects of quantization during training, allowing the model to adapt to low precision representations in advance.

Lesson 5 of 50% complete
โ†Cloud Notebooks: CPU and GPU

Discussion

Sign in to join the discussion

Quantization-Aware Training (QAT)
This chapter first covers GGUF quantized models commonly used in local deployment, helping you understand files, distinguish versions, and make choices. Details on other quantization methods will be covered later.

What Is Inside a GGUF File?

When running models with Ollama, you will frequently encounter names like GGUF, Q4, Q8, etc. These names are mainly related to the model's storage format and quantization method, and are a key reason why large models can run on personal computers.

In August 2023, Georgi Gerganov introduced the GGUF format. Its main advantages include single-file deployment, extensibility, support for memory mapping, complete information, and ease of use. A GGUF file includes a file header, metadata key-value pairs, and tensor information. These components together define the model's structure and behavior. As shown below, GGUF is a model file format for large model inference, currently widely used by local inference tools such as llama.cpp and Ollama. Ollama encapsulates GGUF format handling, model quantization, and model loading processes, so users don't need to convert model formats themselves โ€” they can directly download and run pre-quantized models. A GGUF file contains more than just model parameters; it also includes basic model information, model architecture, and tensor data. GGUF organizes this information through a unified file format, allowing inference tools like llama.cpp to directly read and load it. Its basic structure is shown in the figure below.

ModelScope also supports displaying structured GGUF files, as shown below. When inference tools load GGUF files, they can directly read this information and load the model without needing to prepare multiple model configuration files separately.

GGUF itself is just a model file format and does not directly reduce the model's RAM or VRAM usage. What truly reduces model resource consumption is quantization.

How Should I Read Names Like Q4 and Q8?

In llama.cpp and GGUF models, Q4, Q5, Q8, Q4_K_M represent different quantization names. Here, Q stands for quantization, and the following number indicates the primary quantization bit width. For example, Q4 means primarily 4-bit quantization, and Q8 means primarily 8-bit quantization.

Note that Q4 does not mean all parameters in the model uniformly use 4-bit representation. Taking the common Q4_K_M as an example, it belongs to the K-quant quantization type in llama.cpp. The Q4 in the name means primarily 4-bit quantization, K indicates the use of the K-quant quantization scheme, and M denotes the Medium configuration within that scheme. In practice, different tensors in the model may use different quantization precisions rather than all parameters uniformly using 4-bit representation, as shown in the figure below.

Which Version Should I Download for the Same Model?

When choosing a GGUF model, first confirm that your inference tool supports the model, then look at the specific quantization version. Lower quantization precision generally reduces RAM and VRAM usage, making it easier to run the model on ordinary devices, but may sacrifice some model performance. Higher quantization precision generally preserves more of the original model's capabilities but also requires more hardware resources.

If a version is already close to your device's RAM or VRAM limits, consider switching to a smaller version to leave room for context and runtime operations. If resources are sufficient, you can use the same questions to compare different versions' responses, paying special attention to your target tasks such as information extraction, code generation, or structured output.

First confirm stable operation, then check response quality. This way, the quantized version you choose will be better suited to your device and use case.