AIUnlimited
๐ŸŒณ

AI Foundations

๐ŸŒฑ
AI Seeds

Start from zero

๐ŸŒฟ
AI Sprouts

Build foundations

๐ŸŒณ
AI Branches

Apply in practice

๐Ÿ•๏ธ
AI Canopy

Go deep

๐ŸŒฒ
AI Forest

Master AI

๐Ÿ”จ

AI Mastery

โœ๏ธ
AI Sketch

Start from zero

๐Ÿชจ
AI Chisel

Build foundations

โš’๏ธ
AI Craft

Apply in practice

๐Ÿ’Ž
AI Polish

Go deep

๐Ÿ†
AI Masterpiece

Master AI

๐Ÿ“˜

AI Practice

๐Ÿ“–
Understanding Open-Source Models

Fundamentals and resources for open-source models

๐ŸŽฏ
From Problem to Model Task

Converting business problems to model tasks

โšก
Running Your First Model

See your first results in 30 minutes

๐Ÿ”ง
Fine-Tuning and Evaluation

Fine-tune models and evaluate performance

๐Ÿš€
Application Systems

Build real-world AI applications

๐ŸŽจ
Generative AI

Explore open-source AIGC models

๐Ÿค–
Agents

Learn Agent frameworks and MCP tools

๐Ÿ“
Supplementary Fundamentals

LLM basics and evaluation

๐ŸŽ“

Claude Academy

๐Ÿค–
Claude 101

Learn AI basics with Claude

๐Ÿ’ป
Claude Code 101

Code with Claude as your pair programmer

๐Ÿค
Introduction to Claude Cowork

Collaborate with Claude on complex projects

โš™๏ธ
Claude Platform 101

Build apps with the Claude API

Lab

7 experiments loaded
๐ŸงฌNeural Network Sandbox๐Ÿค–AI or Human?๐Ÿฅ‹Prompt Engineering Dojo๐ŸAlgorithm Race๐Ÿง AI Trivia Challenge๐Ÿ—๏ธSystem Design Canvas
๐ŸŽฏMock InterviewEnter the Labโ†’
๐Ÿš€

Career Development

๐Ÿš€
Interview Launchpad

Start your journey

๐ŸŒŸ
Behavioral Mastery

Master soft skills

๐Ÿ’ป
Technical Interviews

Ace the coding round

๐Ÿค–
AI & ML Interviews

ML interview mastery

๐Ÿ†
Offer & Beyond

Land the best offer

Get Started
AIUnlimited

MIT Licence.

ๆฒชICPๅค‡18025655ๅท-11

Learn

  • AI Basics
  • AI Practice
  • Claude Academy
  • Lab
  • Career Development

Community

  • About
  • FAQ

Support

  • Terms of Service
  • Privacy Policy
  • Contact
AI & Engineering Academicsโ€บโšก Running Your First Modelโ€บLessonsโ€บCloud Notebooks: CPU and GPU
โ˜๏ธ
Running Your First Model โ€ข Beginnerโฑ๏ธ 15 min read

Cloud Notebooks: CPU and GPU

Run Models in the Cloud, Try CPU and GPU with Notebook

The model is running on your laptop, but you want to try a bigger model, or generate an image โ€” and you might not have the right hardware. You can offload computation to the cloud and let the server handle it.

Cloud servers can be configured as needed. We continue using ModelScope Notebook here โ€” select a CPU or GPU runtime instance directly in the browser, download a model and run code. The page runs on your computer; the model runs on the connected cloud instance.

This chapter has two experiments. First, use CPU to run a 0.5B-level model to judge the sentiment of a review. Then use GPU to run an image generation model to turn text descriptions into pictures. Each demonstrates usage under different resources with different tasks and models, so you cannot directly compare CPU vs GPU speed by runtime alone.

Other cloud server resources will be added later.

Open Notebook, Choose the Runtime Instance

From the ModelScope model page, go to "Notebook Quick Development", then click "Connect Runtime". First select a CPU instance for sentiment analysis, then switch to a GPU instance for image generation. For Notebook page operations, refer to the "See Your First Result in 30 Minutes" chapter.

After switching instances, the original Python process and loaded models will not automatically transfer to the new instance. You need to re-check dependencies in the current environment and re-run the experiment code from scratch.

Use CPU to Judge Whether a Review is Positive or Negative

Using the Qwen/Qwen2.5-0.5B-Instruct model as an example, we demonstrate how to use a 0.5B-level model on CPU for sentiment analysis.

After starting the CPU instance, enter the Notebook development environment. Model loading and inference code:

from modelscope import AutoModelForCausalLM, AutoTokenizer
model_name = "Qwen/Qwen2.5-0.5B-Instruct"
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="auto",
    device_map="cpu" 
    # device_map="cpu" # GPU inference
)
print(f"Model device: {model.device}")
tokenizer = AutoTokenizer.from_pretrained(model_name)

# Prompt
prompt = f"""You are a text sentiment classification expert. Please classify the following user review as ใ€Positiveใ€‘, ใ€Negativeใ€‘, or ใ€Neutralใ€‘. Output only one of these three words, without any explanation, punctuation, or extra text.
Review: This movie was really terrible, a waste of time.
Classification result:"""

messages = [
    {"role": "system", "content": "You are Qwen, created by Alibaba Cloud. You are a helpful assistant."},
    {"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)

generated_ids = model.generate(
    **model_inputs,
    max_new_tokens=512
)
generated_ids = [
    output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
]

response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
print("response: ", response )
Lesson 4 of 50% complete
โ†Running Models Locally with Ollama

Discussion

Sign in to join the discussion

Result:

Tutorial illustration

Switch to GPU, Turn Text Descriptions into Images

Using the Tongyi-MAI/Z-Image-Turbo model, we demonstrate how to load a model on GPU and complete an image generation task.

After starting the GPU instance, enter the Notebook development environment.

Per the model card, install the latest diffusers in the terminal:

!pip3 install git+https://github.com/huggingface/diffusers

Installation complete:

Tutorial illustration

Model loading and inference code:

import torch
from modelscope import ZImagePipeline

# 1. Model loading
pipe = ZImagePipeline.from_pretrained(
    "Tongyi-MAI/Z-Image-Turbo",
    torch_dtype=torch.bfloat16,
    low_cpu_mem_usage=False,
)
pipe.to("cuda")
print(f"Model device: {pipe.device}")

prompt = "A young Chinese woman dressed in a red hanfu, its sleeves and hem embroidered with intricate, elaborate patterns. Her makeup is flawless, with a red huadian mark on her forehead adding a touch of classical elegance. Her towering hairstyle is majestic and dignified, topped with a golden phoenix crown and adorned with red flowers and pearl strings that sway gracefully. She holds a round folding fan painted with a lady, an ancient tree, and birds in a tranquil mood. A neon lightning-shaped lamp hangs above her outstretched left palm, casting a bright warm-yellow glow. The night is gentle, the lantern light hazy; in the distance, the layered tiers of Xi'an's Giant Wild Goose Pagoda emerge faintly from a soft halo, with dappled bokeh in the background, dreamlike and ethereal."
# 2. Generate image
image = pipe(
    prompt=prompt,
    height=1024,
    width=1024,
    num_inference_steps=9, 
    guidance_scale=0.0, 
    generator=torch.Generator("cuda").manual_seed(42),
).images[0]
print("Image generated")

image.save("image-vintage.png")
print("Image saved successfully!")

Result:

Tutorial illustration

Generated image:

Tutorial illustration

Code and related files for this chapter: https://www.modelscope.cn/gallery/liucong/4a8eb492-54dc-4062-82c0-65b17fd94d5f