After running your first result, it's easy to want to try larger models. You can switch from 0.5B to 7B, but the original machine may no longer be sufficient.
When selecting resources, first distinguish between a few things. CPU and GPU handle computation, memory and video memory hold data during execution, and disk stores model files. A model that downloads successfully only means the disk has enough space โ whether it can run smoothly also depends on memory, video memory, and computing power. The same model, depending on what precision you load it at, how long the input is, and how many requests are processed simultaneously, will all affect resource usage. Inference and training also have different requirements.
Let's first look at what CPU and GPU are each suited for, then estimate how much space model parameters will occupy, to prepare for choosing a laptop or cloud instance later.
CPU and GPU differ significantly in their computing approaches.
CPU is designed for general-purpose computing, with each core having relatively complete control and computation capabilities. It can make decisions, jump, and schedule tasks based on the program execution process, flexibly handling different types of computing tasks.
GPU uses a parallel computing architecture, containing a large number of compute units that can execute a large number of computing tasks simultaneously. It is suited for handling tasks with large data scales and high computational repetitiveness.
CPU and GPU processing speed depends on the task type and computation scale.
For scenarios involving a large amount of logical judgment, program control, and task scheduling, CPU has higher execution efficiency. Tasks like operating system operation and database queries are well-suited for CPU.
For large-scale parallel computing tasks, GPU can provide higher computational throughput. Matrix multiplication and convolution operations in deep learning can be split into multiple parallel subtasks, executed simultaneously by GPU compute units, improving overall computing efficiency.
However, the advantages of GPU only become significant when the scale of computing tasks is large. When the data volume is small, GPU computing resources cannot be fully utilized, and data transfer and task scheduling between CPU and GPU also bring additional overhead. In this case, using CPU may be more efficient.
The costs of CPU and GPU not only include hardware purchase prices, but also power consumption during operation, as well as subsequent usage and maintenance costs.
In terms of hardware procurement, CPU is standard โ ordinary servers and personal computers all need CPU configuration. High-performance GPUs used for AI training and inference are typically expensive. Besides the GPU itself, you also need to consider the costs of servers, video memory, storage, and other supporting hardware.
Sign in to join the discussion
During operation, GPU has high power consumption when performing large-scale computations. The more GPUs in a server, the greater the pressure on power supply and cooling, and the electricity and cooling costs of data centers will increase significantly.
In terms of usage and maintenance, the CPU ecosystem is relatively mature, and routine server maintenance is relatively straightforward. GPU requires consideration of driver, computing framework, and hardware environment compatibility. Conflicts between different versions may prevent models from running normally. Multi-GPU environments also involve data communication between devices, which has certain requirements for hardware connections and software configuration.
Training and inference tasks have different resource requirements, and models of different scales and types also have significant differences in computing resources. Hardware selection needs to comprehensively consider model scale, data volume, performance requirements, and resource needs in both training and inference scenarios.
In the model training phase, when the data scale is small, CPU can meet the training needs of most traditional machine learning models, such as linear regression, logistic regression, and random forests. Some lightweight neural networks can also use CPU for training, but as data scale and model structure complexity increase, training time will increase significantly. If you need to frequently tune models, using GPU can reduce training time.
Large-scale neural networks and large language models require extensive matrix computation for training, with high requirements for computing power and video memory capacity. Large language models with billions of parameters, like Qwen, need to rely on high-performance acceleration devices like GPU for training.
Model requirements during inference differ from training. For scenarios with smaller parameter scales and lower access volumes, CPU can meet inference needs, and resource usage can also be reduced through methods like quantization. For large models, high-concurrency services, and scenarios requiring fast response, GPU can provide higher inference efficiency.
CPU and GPU each have applicable scenarios, and you need to comprehensively consider model scale, data volume, concurrent requests, performance requirements, and deployment costs.
Evaluating a model's resource requirements mainly depends on two indicators: parameter scale and data precision.
Model parameter scale is generally expressed in B, where 1B represents 1 billion parameters. Parameters are the numerical values the model learns during training. During inference, these parameters need to be loaded into memory (CPU) or video memory (GPU). The larger the parameter count, the more resources occupied.
Data precision refers to what numerical format the model parameters use for storage and computation. Common formats include FP32, FP16, BF16, and quantization formats like INT8 and INT4. For the same parameter count, higher data precision means more resources occupied.
The estimated storage space for 1B parameters at different precisions is as follows:
| Precision | Space per 1B Parameters (approx.) | Common Uses |
| FP32 (Single Precision) | 4GB | Training, high-precision inference |
| FP16 (Half Precision) | 2GB | Training, inference |
| BF16 | 2GB | Training, inference |
| INT8 | 1GB | Quantized inference |
| INT4 | 0.5GB | Low-precision quantized inference |
Model weights = Parameter count ร Storage space per 1B parameters at that precision. For a 7B model with FP16 precision inference, the model weight occupation is approximately: 7 ร 2GB โ 14GB.
During model inference, besides loading model weights, additional resources are needed for KV Cache, runtime framework, and temporary computation.
KV Cache is used to store Token-related information that has already been calculated during the generation process, avoiding redundant computation. Its usage depends on the model structure, context length, and number of concurrent requests. The longer the context, the more information needs to be stored; the more requests processed simultaneously, the more KV Cache usage increases.
Estimation formula for inference resource usage:
CPU inference required memory โ Model weights + Framework and temporary overhead.
GPU inference required video memory โ Model weights + KV Cache + Framework and temporary overhead.
Among these, weights are static, framework overhead is basically constant, and the variable is entirely in KV Cache. In long sequence scenarios, KV Cache space needs to be reserved.
Model training phase resource requirements are higher than inference phase. Besides storing model weights, gradients, optimizer states, and intermediate results during training also need to be stored, requiring more resources. Training resource requirements depend on optimizer type and training strategy.
(1) Full Model Training
For full model training, model weights, gradients, optimizer states, and activations need to be stored simultaneously. Taking full training with Adam optimizer and FP16 precision as an example, each 1B parameter requires approximately:
| Component | Per 1B Parameters |
| Model weights (FP16) | 2 GB |
| Gradients (FP16) | 2 GB |
| Optimizer states | 8 GB |
| FP32 weight copy | 4 GB |
| Subtotal (excluding activations) | 16 GB |
Note: The Adam optimizer needs to save two copies of FP32 precision states (momentum and variance), so optimizer states take 8GB; the FP32 weight copy is used for optimizer parameter updates.
The preliminary estimation formula for full training is:
GPU training required video memory โ Parameter count ร 16 Byte + Activation overhead
CPU training required memory โ Parameter count ร 16 Byte + Activation overhead
GPU training is limited by video memory capacity, while CPU training is limited by computation speed. For example, a 7B model with FP16 full training on GPU requires at least 7B ร 16 Byte โ 112GB of video memory, and the actual requirement will be larger after adding activations.
(2) Parameter-Efficient Fine-Tuning (e.g., LoRA)
For large model fine-tuning, parameter-efficient fine-tuning methods like LoRA (Low-Rank Adaptation) are more commonly used in practice. LoRA does not update all original model parameters during training; instead, it freezes the original model weights and only trains newly added low-rank adaptation parameters, significantly reducing video memory requirements.
LoRA training video memory โ Base model weights + LoRA parameters + LoRA parameter gradients and optimizer states + Activations.
Among these, base model weights only need to be loaded and do not participate in updates; LoRA added parameters are only 0.1%~1% of the original model, and their gradients and optimizer overhead are almost negligible. Actual video memory mainly depends on model weights, context length, batch size, and training framework optimization methods. LoRA fine-tuning can significantly reduce video memory requirements compared to full training.
For example, loading a 7B model with FP16, model weights take 4GB. On a single GPU with 24GB video memory, LoRA fine-tuning can usually run (with appropriate adjustments to batch size and context length). Using QLoRA (quantizing the base model to INT4) can further reduce video memory usage, enabling fine-tuning on more limited hardware.