AIUnlimited
๐ŸŒณ

AI Foundations

๐ŸŒฑ
AI Seeds

Start from zero

๐ŸŒฟ
AI Sprouts

Build foundations

๐ŸŒณ
AI Branches

Apply in practice

๐Ÿ•๏ธ
AI Canopy

Go deep

๐ŸŒฒ
AI Forest

Master AI

๐Ÿ”จ

AI Mastery

โœ๏ธ
AI Sketch

Start from zero

๐Ÿชจ
AI Chisel

Build foundations

โš’๏ธ
AI Craft

Apply in practice

๐Ÿ’Ž
AI Polish

Go deep

๐Ÿ†
AI Masterpiece

Master AI

๐Ÿ“˜

AI Practice

๐Ÿ“–
Understanding Open-Source Models

Fundamentals and resources for open-source models

๐ŸŽฏ
From Problem to Model Task

Converting business problems to model tasks

โšก
Running Your First Model

See your first results in 30 minutes

๐Ÿ”ง
Fine-Tuning and Evaluation

Fine-tune models and evaluate performance

๐Ÿš€
Application Systems

Build real-world AI applications

๐ŸŽจ
Generative AI

Explore open-source AIGC models

๐Ÿค–
Agents

Learn Agent frameworks and MCP tools

๐Ÿ“
Supplementary Fundamentals

LLM basics and evaluation

๐ŸŽ“

Claude Academy

๐Ÿค–
Claude 101

Learn AI basics with Claude

๐Ÿ’ป
Claude Code 101

Code with Claude as your pair programmer

๐Ÿค
Introduction to Claude Cowork

Collaborate with Claude on complex projects

โš™๏ธ
Claude Platform 101

Build apps with the Claude API

Lab

7 experiments loaded
๐ŸงฌNeural Network Sandbox๐Ÿค–AI or Human?๐Ÿฅ‹Prompt Engineering Dojo๐ŸAlgorithm Race๐Ÿง AI Trivia Challenge๐Ÿ—๏ธSystem Design Canvas
๐ŸŽฏMock InterviewEnter the Labโ†’
๐Ÿš€

Career Development

๐Ÿš€
Interview Launchpad

Start your journey

๐ŸŒŸ
Behavioral Mastery

Master soft skills

๐Ÿ’ป
Technical Interviews

Ace the coding round

๐Ÿค–
AI & ML Interviews

ML interview mastery

๐Ÿ†
Offer & Beyond

Land the best offer

Get Started
AIUnlimited

MIT Licence.

ๆฒชICPๅค‡18025655ๅท-11

Learn

  • AI Basics
  • AI Practice
  • Claude Academy
  • Lab
  • Career Development

Community

  • About
  • FAQ

Support

  • Terms of Service
  • Privacy Policy
  • Contact
AI & Engineering Academicsโ€บโšก Running Your First Modelโ€บLessonsโ€บServer Configuration for Models
๐Ÿ–ฅ๏ธ
Running Your First Model โ€ข Beginnerโฑ๏ธ 15 min read

Server Configuration for Models

Choosing the Right Server Configuration for Your Model Size

After running your first result, it's easy to want to try larger models. You can switch from 0.5B to 7B, but the original machine may no longer be sufficient.

When selecting resources, first distinguish between a few things. CPU and GPU handle computation, memory and video memory hold data during execution, and disk stores model files. A model that downloads successfully only means the disk has enough space โ€” whether it can run smoothly also depends on memory, video memory, and computing power. The same model, depending on what precision you load it at, how long the input is, and how many requests are processed simultaneously, will all affect resource usage. Inference and training also have different requirements.

Let's first look at what CPU and GPU are each suited for, then estimate how much space model parameters will occupy, to prepare for choosing a laptop or cloud instance later.

CPU and GPU: What Is Each Good At?

One Excels at Flexible Processing, the Other at Parallel Computing

CPU and GPU differ significantly in their computing approaches.

CPU is designed for general-purpose computing, with each core having relatively complete control and computation capabilities. It can make decisions, jump, and schedule tasks based on the program execution process, flexibly handling different types of computing tasks.

GPU uses a parallel computing architecture, containing a large number of compute units that can execute a large number of computing tasks simultaneously. It is suited for handling tasks with large data scales and high computational repetitiveness.

Is GPU Faster at Everything?

CPU and GPU processing speed depends on the task type and computation scale.

For scenarios involving a large amount of logical judgment, program control, and task scheduling, CPU has higher execution efficiency. Tasks like operating system operation and database queries are well-suited for CPU.

For large-scale parallel computing tasks, GPU can provide higher computational throughput. Matrix multiplication and convolution operations in deep learning can be split into multiple parallel subtasks, executed simultaneously by GPU compute units, improving overall computing efficiency.

However, the advantages of GPU only become significant when the scale of computing tasks is large. When the data volume is small, GPU computing resources cannot be fully utilized, and data transfer and task scheduling between CPU and GPU also bring additional overhead. In this case, using CPU may be more efficient.

Besides Buying Equipment, What Other Costs Should You Consider?

The costs of CPU and GPU not only include hardware purchase prices, but also power consumption during operation, as well as subsequent usage and maintenance costs.

In terms of hardware procurement, CPU is standard โ€” ordinary servers and personal computers all need CPU configuration. High-performance GPUs used for AI training and inference are typically expensive. Besides the GPU itself, you also need to consider the costs of servers, video memory, storage, and other supporting hardware.

Lesson 2 of 50% complete
โ†Your First Inference in 30 Minutes

Discussion

Sign in to join the discussion

During operation, GPU has high power consumption when performing large-scale computations. The more GPUs in a server, the greater the pressure on power supply and cooling, and the electricity and cooling costs of data centers will increase significantly.

In terms of usage and maintenance, the CPU ecosystem is relatively mature, and routine server maintenance is relatively straightforward. GPU requires consideration of driver, computing framework, and hardware environment compatibility. Conflicts between different versions may prevent models from running normally. Multi-GPU environments also involve data communication between devices, which has certain requirements for hardware connections and software configuration.

How to Choose Between CPU and GPU?

Training and inference tasks have different resource requirements, and models of different scales and types also have significant differences in computing resources. Hardware selection needs to comprehensively consider model scale, data volume, performance requirements, and resource needs in both training and inference scenarios.

In the model training phase, when the data scale is small, CPU can meet the training needs of most traditional machine learning models, such as linear regression, logistic regression, and random forests. Some lightweight neural networks can also use CPU for training, but as data scale and model structure complexity increase, training time will increase significantly. If you need to frequently tune models, using GPU can reduce training time.

Large-scale neural networks and large language models require extensive matrix computation for training, with high requirements for computing power and video memory capacity. Large language models with billions of parameters, like Qwen, need to rely on high-performance acceleration devices like GPU for training.

Model requirements during inference differ from training. For scenarios with smaller parameter scales and lower access volumes, CPU can meet inference needs, and resource usage can also be reduced through methods like quantization. For large models, high-concurrency services, and scenarios requiring fast response, GPU can provide higher inference efficiency.

CPU and GPU each have applicable scenarios, and you need to comprehensively consider model scale, data volume, concurrent requests, performance requirements, and deployment costs.

How Big Is the Model? How Much Memory and Video Memory Do You Need?

First Look at Parameter Count, Then Loading Precision

Evaluating a model's resource requirements mainly depends on two indicators: parameter scale and data precision.

Model parameter scale is generally expressed in B, where 1B represents 1 billion parameters. Parameters are the numerical values the model learns during training. During inference, these parameters need to be loaded into memory (CPU) or video memory (GPU). The larger the parameter count, the more resources occupied.

Data precision refers to what numerical format the model parameters use for storage and computation. Common formats include FP32, FP16, BF16, and quantization formats like INT8 and INT4. For the same parameter count, higher data precision means more resources occupied.

The estimated storage space for 1B parameters at different precisions is as follows:

PrecisionSpace per 1B Parameters (approx.)Common Uses
FP32 (Single Precision) 4GBTraining, high-precision inference
FP16 (Half Precision) 2GBTraining, inference
BF16 2GBTraining, inference
INT8 1GBQuantized inference
INT4 0.5GBLow-precision quantized inference

Model weights = Parameter count ร— Storage space per 1B parameters at that precision. For a 7B model with FP16 precision inference, the model weight occupation is approximately: 7 ร— 2GB โ‰ˆ 14GB.

First Look at Parameter Count, Then Loading Precision

During model inference, besides loading model weights, additional resources are needed for KV Cache, runtime framework, and temporary computation.

KV Cache is used to store Token-related information that has already been calculated during the generation process, avoiding redundant computation. Its usage depends on the model structure, context length, and number of concurrent requests. The longer the context, the more information needs to be stored; the more requests processed simultaneously, the more KV Cache usage increases.

Estimation formula for inference resource usage:

CPU inference required memory โ‰ˆ Model weights + Framework and temporary overhead.

GPU inference required video memory โ‰ˆ Model weights + KV Cache + Framework and temporary overhead.

Among these, weights are static, framework overhead is basically constant, and the variable is entirely in KV Cache. In long sequence scenarios, KV Cache space needs to be reserved.

For the Same Model, Why Does Training Use More Resources?

Model training phase resource requirements are higher than inference phase. Besides storing model weights, gradients, optimizer states, and intermediate results during training also need to be stored, requiring more resources. Training resource requirements depend on optimizer type and training strategy.

(1) Full Model Training

For full model training, model weights, gradients, optimizer states, and activations need to be stored simultaneously. Taking full training with Adam optimizer and FP16 precision as an example, each 1B parameter requires approximately:

ComponentPer 1B Parameters
Model weights (FP16)2 GB
Gradients (FP16)2 GB
Optimizer states8 GB
FP32 weight copy4 GB
Subtotal (excluding activations)16 GB

Note: The Adam optimizer needs to save two copies of FP32 precision states (momentum and variance), so optimizer states take 8GB; the FP32 weight copy is used for optimizer parameter updates.

The preliminary estimation formula for full training is:

GPU training required video memory โ‰ˆ Parameter count ร— 16 Byte + Activation overhead

CPU training required memory โ‰ˆ Parameter count ร— 16 Byte + Activation overhead

GPU training is limited by video memory capacity, while CPU training is limited by computation speed. For example, a 7B model with FP16 full training on GPU requires at least 7B ร— 16 Byte โ‰ˆ 112GB of video memory, and the actual requirement will be larger after adding activations.

(2) Parameter-Efficient Fine-Tuning (e.g., LoRA)

For large model fine-tuning, parameter-efficient fine-tuning methods like LoRA (Low-Rank Adaptation) are more commonly used in practice. LoRA does not update all original model parameters during training; instead, it freezes the original model weights and only trains newly added low-rank adaptation parameters, significantly reducing video memory requirements.

LoRA training video memory โ‰ˆ Base model weights + LoRA parameters + LoRA parameter gradients and optimizer states + Activations.

Among these, base model weights only need to be loaded and do not participate in updates; LoRA added parameters are only 0.1%~1% of the original model, and their gradients and optimizer overhead are almost negligible. Actual video memory mainly depends on model weights, context length, batch size, and training framework optimization methods. LoRA fine-tuning can significantly reduce video memory requirements compared to full training.

For example, loading a 7B model with FP16, model weights take 4GB. On a single GPU with 24GB video memory, LoRA fine-tuning can usually run (with appropriate adjustments to batch size and context length). Using QLoRA (quantizing the base model to INT4) can further reduce video memory usage, enabling fine-tuning on more limited hardware.