AIUnlimited
๐ŸŒณ

AI Foundations

๐ŸŒฑ
AI Seeds

Start from zero

๐ŸŒฟ
AI Sprouts

Build foundations

๐ŸŒณ
AI Branches

Apply in practice

๐Ÿ•๏ธ
AI Canopy

Go deep

๐ŸŒฒ
AI Forest

Master AI

๐Ÿ”จ

AI Mastery

โœ๏ธ
AI Sketch

Start from zero

๐Ÿชจ
AI Chisel

Build foundations

โš’๏ธ
AI Craft

Apply in practice

๐Ÿ’Ž
AI Polish

Go deep

๐Ÿ†
AI Masterpiece

Master AI

๐Ÿ“˜

AI Practice

๐Ÿ“–
Understanding Open-Source Models

Fundamentals and resources for open-source models

๐ŸŽฏ
From Problem to Model Task

Converting business problems to model tasks

โšก
Running Your First Model

See your first results in 30 minutes

๐Ÿ”ง
Fine-Tuning and Evaluation

Fine-tune models and evaluate performance

๐Ÿš€
Application Systems

Build real-world AI applications

๐ŸŽจ
Generative AI

Explore open-source AIGC models

๐Ÿค–
Agents

Learn Agent frameworks and MCP tools

๐Ÿ“
Supplementary Fundamentals

LLM basics and evaluation

๐ŸŽ“

Claude Academy

๐Ÿค–
Claude 101

Learn AI basics with Claude

๐Ÿ’ป
Claude Code 101

Code with Claude as your pair programmer

๐Ÿค
Introduction to Claude Cowork

Collaborate with Claude on complex projects

โš™๏ธ
Claude Platform 101

Build apps with the Claude API

Lab

7 experiments loaded
๐ŸงฌNeural Network Sandbox๐Ÿค–AI or Human?๐Ÿฅ‹Prompt Engineering Dojo๐ŸAlgorithm Race๐Ÿง AI Trivia Challenge๐Ÿ—๏ธSystem Design Canvas
๐ŸŽฏMock InterviewEnter the Labโ†’
๐Ÿš€

Career Development

๐Ÿš€
Interview Launchpad

Start your journey

๐ŸŒŸ
Behavioral Mastery

Master soft skills

๐Ÿ’ป
Technical Interviews

Ace the coding round

๐Ÿค–
AI & ML Interviews

ML interview mastery

๐Ÿ†
Offer & Beyond

Land the best offer

Get Started
AIUnlimited

MIT Licence.

ๆฒชICPๅค‡18025655ๅท-11

Learn

  • AI Basics
  • AI Practice
  • Claude Academy
  • Lab
  • Career Development

Community

  • About
  • FAQ

Support

  • Terms of Service
  • Privacy Policy
  • Contact
AI & Engineering Academicsโ€บ๐Ÿ“– Understanding Open-Source Modelsโ€บLessonsโ€บData: The Foundation of Fine-Tuning
๐Ÿ“Š
Understanding Open-Source Models โ€ข Beginnerโฑ๏ธ 12 min read

Data: The Foundation of Fine-Tuning

Data: The Foundation of Open Source Model Fine-Tuning

Have you ever encountered a situation where a model excels at coding and math, but gives random answers to everyday questions?

We all know that during training, to improve a specific capability, you need to find that type of data,

But many datasets are hard to find through direct internet search. A unified download portal saves enormous effort for model trainers.

For example, our open-source Chinese DeepSeek-R1 distilled dataset had much of its raw data downloaded directly from ModelScope,

illustration

Fine-tuning can start with an existing open-source dataset, allowing you to quickly run through the entire pipeline.

Finding Data Requires Knowing What You Plan to Do

ModelScope has over 40,000 open-source datasets, supporting search and preview with field descriptions and versions.

ModelScope supports searching datasets by keyword tags and task categories, such as searching for Chinese binary text classification data,

illustration

Before Downloading, Preview a Few Samples

After filtering, select a dataset and click 'Data Preview' to view samples and field descriptions, such as simpleai/HC3-Chinese.

illustration

Which column has the question, which has the answer, whether there are labels, and whether each record is text or dialogue -- all can be checked upfront.

For QA data, human and model answers may be in different fields. Decide which fields to read and how to organize before training.

Check dataset files and versions. Datasets get updated; record which version you used for reproducibility.

illustration

Dataset descriptions and licenses are worth reviewing. Downloaded data may have different requirements for research, training, or commercial use.

Load the Data First, Then Decide What to Use

ModelScope provides the MsDataset.load interface to load datasets in Python. After installing ModelScope, follow the dataset page's instructions.

Load a Complete Dataset First

Using DAMO_NLP/jd as an example, read the default configuration like this,

from modelscope.msdatasets import MsDataset

ds = MsDataset.load(
    'DAMO_NLP/jd',
    trust_remote_code=True,
)
print(ds)
Lesson 3 of 40% complete
โ†Model Discovery and Downloads

Discussion

Sign in to join the discussion

illustration

If You Only Need a Subset, Specify the Name Clearly

Some datasets separate data by source, domain, or purpose. If you only need part, use subset_name to specify the subset.

For example, load the baike subset from HC3-Chinese.

from modelscope.msdatasets import MsDataset

ds = MsDataset.load(
    'simpleai/HC3-Chinese',
    subset_name='baike',
    trust_remote_code=True,
)
print(ds)
illustration

When switching datasets, check available subsets. baike is specific to this dataset.

Training and Test Sets Can Be Loaded Separately

The subset selects the data category, while split selects the partition. Common splits are train, validation, and test.

To load only the training set from baike, add another parameter.

from modelscope.msdatasets import MsDataset

ds = MsDataset.load(
    'simpleai/HC3-Chinese',
    subset_name='baike',
    split='train',
    trust_remote_code=True,
)
print(ds[0])
illustration

Not all datasets provide all three splits. Print a sample to verify fields before proceeding.

To Reproduce Experiments, Record the Version Too

After data updates, the same code may read different content. Using HC3-Chinese v1 as an example, specify via version.

from modelscope.msdatasets import MsDataset

ds = MsDataset.load(
    'simpleai/HC3-Chinese',
    subset_name='baike',
    version='v1',
    trust_remote_code=True,
)
print(ds)
illustration

Check available versions on the dataset page. If a version updates continuously, keep the files or checksums you used.

Data Downloaded? Don't Rush to Feed It All to the Model

First check a few samples for null values, garbled text, or errors, then verify fields meet training requirements.

Training and test data must be kept separate. Don't mix test samples into training sets.