AIUnlimited
๐ŸŒณ

AI Foundations

๐ŸŒฑ
AI Seeds

Start from zero

๐ŸŒฟ
AI Sprouts

Build foundations

๐ŸŒณ
AI Branches

Apply in practice

๐Ÿ•๏ธ
AI Canopy

Go deep

๐ŸŒฒ
AI Forest

Master AI

๐Ÿ”จ

AI Mastery

โœ๏ธ
AI Sketch

Start from zero

๐Ÿชจ
AI Chisel

Build foundations

โš’๏ธ
AI Craft

Apply in practice

๐Ÿ’Ž
AI Polish

Go deep

๐Ÿ†
AI Masterpiece

Master AI

๐Ÿ“˜

AI Practice

๐Ÿ“–
Understanding Open-Source Models

Fundamentals and resources for open-source models

๐ŸŽฏ
From Problem to Model Task

Converting business problems to model tasks

โšก
Running Your First Model

See your first results in 30 minutes

๐Ÿ”ง
Fine-Tuning and Evaluation

Fine-tune models and evaluate performance

๐Ÿš€
Application Systems

Build real-world AI applications

๐ŸŽจ
Generative AI

Explore open-source AIGC models

๐Ÿค–
Agents

Learn Agent frameworks and MCP tools

๐Ÿ“
Supplementary Fundamentals

LLM basics and evaluation

๐ŸŽ“

Claude Academy

๐Ÿค–
Claude 101

Learn AI basics with Claude

๐Ÿ’ป
Claude Code 101

Code with Claude as your pair programmer

๐Ÿค
Introduction to Claude Cowork

Collaborate with Claude on complex projects

โš™๏ธ
Claude Platform 101

Build apps with the Claude API

Lab

7 experiments loaded
๐ŸงฌNeural Network Sandbox๐Ÿค–AI or Human?๐Ÿฅ‹Prompt Engineering Dojo๐ŸAlgorithm Race๐Ÿง AI Trivia Challenge๐Ÿ—๏ธSystem Design Canvas
๐ŸŽฏMock InterviewEnter the Labโ†’
๐Ÿš€

Career Development

๐Ÿš€
Interview Launchpad

Start your journey

๐ŸŒŸ
Behavioral Mastery

Master soft skills

๐Ÿ’ป
Technical Interviews

Ace the coding round

๐Ÿค–
AI & ML Interviews

ML interview mastery

๐Ÿ†
Offer & Beyond

Land the best offer

Get Started
AIUnlimited

MIT Licence.

ๆฒชICPๅค‡18025655ๅท-11

Learn

  • AI Basics
  • AI Practice
  • Claude Academy
  • Lab
  • Career Development

Community

  • About
  • FAQ

Support

  • Terms of Service
  • Privacy Policy
  • Contact
AI & Engineering Academicsโ€บโšก Running Your First Modelโ€บLessonsโ€บRunning Models Locally with Ollama
๐Ÿ’ป
Running Your First Model โ€ข Beginnerโฑ๏ธ 20 min read

Running Models Locally with Ollama

Run Open Source Models on Your Laptop, Start with Ollama

Models can be deployed locally or in the cloud. Local devices include laptops, desktops, and other edge devices, etc.

There are also different tool choices for local running: Ollama, as well as frameworks like vLLM and SGLang.

In this chapter, we'll first explain the basic usage of Ollama clearly. Adaptation for other edge devices and the use of frameworks like vLLM and SGLang will be covered later.

Why Would You Want to Run Models on Your Own Computer?

Calling large models through an API is the most hassle-free approach: send the request, the model calculates and returns the result. Users don't need to purchase GPUs or maintain the underlying environment.

But the API approach also has unavoidable limitations, mainly in the following aspects:

  • Data security: core documents, customer privacy, and internal business data cannot be directly sent to third-party public clouds.
  • Offline availability: production environments like intranet dedicated lines simply cannot connect to the external network.
  • Cost control: when call volume is large and stable, building your own server is usually more cost-effective than paying per token.

So, if your business involves data sensitivity, requires purely offline operation, or has large call volumes, local deployment is the more appropriate choice.

First Check Your Laptop's System and Hardware

Ollama is one of the currently more user-friendly tools, packaging model downloading, local management, and API services all together. Users don't need to write code to load models; just a few commands can start the model directly.

Whether a model can run smoothly locally largely depends on the computer's hardware resources. Ollama provides good support for common hardware platforms, supporting Windows, Linux, and macOS. Whether it's a regular CPU, GPU, or a Mac with Apple Silicon, they can all be used to run models.

CPU running, even without a dedicated graphics card, you can run models through CPU. Since large model inference requires extensive computation, generation speed is usually slower with CPU, making it more suitable for smaller parameter models or scenarios where response speed requirements are not high.

GPU running, this is currently the more common way to run large models. GPU has strong parallel computing capabilities that can significantly improve model generation speed. The larger the model, the more video memory is usually needed, so when choosing a model, you need to consider whether the graphics card has enough video memory.

Apple Silicon running, Apple's M-series chips use unified memory design, where CPU and GPU can use the same memory. For Mac users, if the device has 32GB, 64GB or larger unified memory, you can try running some larger parameter quantized models.

Lesson 3 of 50% complete
โ†Server Configuration for Models

Discussion

Sign in to join the discussion

Choose Your Operating System, Install Ollama

Install Ollama on Windows

  1. First, on the Ollama official website, Download Ollama and download the corresponding Windows version
illustration
  1. After download is complete, click install directly, click continue
illustration
  1. After installation is complete, it shows that the Ollama service has started
illustration
  1. In cmd, start the model you want to load, for example: qwen3.5:2b, type the command directly in the terminal
ollama run qwen3.5:2b

The page output shows that the model has started successfully. At this point you can have conversations directly in the client or call the API to test the model.

illustration

Test API code:

from openai import OpenAI

client = OpenAI(
    base_url="http://127.0.0.1:11434/v1",
    api_key="ollama",  
)

response = client.chat.completions.create(
    model="qwen3.5:2b", 
    messages=[
        {
            "role": "system",
            "content": "You are a helpful assistant.",
        },
        {
            "role": "user",
            "content": "Who are you?",
        },
    ],
    temperature=0,
    max_tokens=512,
)

print(response.choices[0].message.content)

Result:

illustration

Install Ollama on ModelScope Notebook

  1. Install Ollama in the Notebook, first download the installation package, execute the following command
!modelscope download --model=modelscope/ollama-linux --local_dir ./ollama-linux  
illustration
  1. After the installation package is downloaded, enter the ollama-linux folder and execute the following command to install
%cd ollama-linux
illustration
sudo chmod 777 ./ollama-modelscope-install.sh
./ollama-modelscope-install.sh

After installation is complete, it will show API endpoint information, default 127.0.0.1:11434. If the following information is shown, the installation is complete.

illustration

Install Ollama on Linux

The installation steps on Linux are the same as installing Ollama in ModelScope Notebook above.

Install Ollama on Mac

  1. First, on the Ollama official website, Download Ollama and download the corresponding macOS version
illustration
  1. After download is complete, double-click the downloaded file and drag ollama into Applications
illustration
  1. After installation is complete, it shows that the Ollama service has started
illustration
  1. In cmd, start the model you want to load, for example: qwen3.5:2b, type the command directly in the terminal
ollama run qwen3.5:2b

The page output shows that the model has started successfully. At this point you can have conversations directly in the client or call the API to test the model.

illustration

You can also chat in the client.

illustration

You Can Also Download Models from ModelScope and Start Them with Ollama

Ollama provides many model repositories. We can find the models we need in ModelScope and start them with Ollama.

  1. Start the Ollama service

For long-running services, Jupyter's single kernel can only execute one cell at a time, making it difficult to execute subsequent cells. So this part should be executed in the terminal. After opening the Notebook, go directly to the workspace, where you'll see "Terminal" at the bottom. Click it.

illustration

Enter in the terminal:

ollama serve
illustration
  1. Next, run the model we want to start, for example: Qwen3-4B-GGUF. The first time the model runs, it needs to download the corresponding files from the model repository. This command also needs to be executed in the terminal. You can open a new terminal and execute the following command:
illustration
 ollama run modelscope.cn/unsloth/Qwen3-4B-GGUF
illustration

At this point, the model has started. You can call the API to test whether the model endpoint is connected. The code is as follows:

import requests
url = "http://127.0.0.1:11434/v1/chat/completions"
data = {
    "model": "modelscope.cn/unsloth/Qwen3-4B-GGUF",
    "messages": [
        {
            "role": "user",
            "content": "Hello, please introduce yourself"
        }
    ],
    "max_tokens": 100
}

response = requests.post(url, json=data)

print("Status code:", response.status_code)
print(response.text)

If the model call is successful, it will output the corresponding model result.

illustration

The complete experimental process involved in this chapter can be found at: https://modelscope.cn/gallery/liucong/b1ef397b-fe80-4fb9-812f-969fa25ea44c