Have you ever encountered a situation where a model excels at coding and math, but gives random answers to everyday questions?
We all know that during training, to improve a specific capability, you need to find that type of data,
But many datasets are hard to find through direct internet search. A unified download portal saves enormous effort for model trainers.
For example, our open-source Chinese DeepSeek-R1 distilled dataset had much of its raw data downloaded directly from ModelScope,
Fine-tuning can start with an existing open-source dataset, allowing you to quickly run through the entire pipeline.
ModelScope has over 40,000 open-source datasets, supporting search and preview with field descriptions and versions.
ModelScope supports searching datasets by keyword tags and task categories, such as searching for Chinese binary text classification data,
After filtering, select a dataset and click 'Data Preview' to view samples and field descriptions, such as simpleai/HC3-Chinese.
Which column has the question, which has the answer, whether there are labels, and whether each record is text or dialogue -- all can be checked upfront.
For QA data, human and model answers may be in different fields. Decide which fields to read and how to organize before training.
Check dataset files and versions. Datasets get updated; record which version you used for reproducibility.
Dataset descriptions and licenses are worth reviewing. Downloaded data may have different requirements for research, training, or commercial use.
ModelScope provides the MsDataset.load interface to load datasets in Python. After installing ModelScope, follow the dataset page's instructions.
Using DAMO_NLP/jd as an example, read the default configuration like this,
from modelscope.msdatasets import MsDataset
ds = MsDataset.load(
'DAMO_NLP/jd',
trust_remote_code=True,
)
print(ds)
Sign in to join the discussion

Some datasets separate data by source, domain, or purpose. If you only need part, use subset_name to specify the subset.
For example, load the baike subset from HC3-Chinese.
from modelscope.msdatasets import MsDataset
ds = MsDataset.load(
'simpleai/HC3-Chinese',
subset_name='baike',
trust_remote_code=True,
)
print(ds)

When switching datasets, check available subsets. baike is specific to this dataset.
The subset selects the data category, while split selects the partition. Common splits are train, validation, and test.
To load only the training set from baike, add another parameter.
from modelscope.msdatasets import MsDataset
ds = MsDataset.load(
'simpleai/HC3-Chinese',
subset_name='baike',
split='train',
trust_remote_code=True,
)
print(ds[0])

Not all datasets provide all three splits. Print a sample to verify fields before proceeding.
After data updates, the same code may read different content. Using HC3-Chinese v1 as an example, specify via version.
from modelscope.msdatasets import MsDataset
ds = MsDataset.load(
'simpleai/HC3-Chinese',
subset_name='baike',
version='v1',
trust_remote_code=True,
)
print(ds)

Check available versions on the dataset page. If a version updates continuously, keep the files or checksums you used.
First check a few samples for null values, garbled text, or errors, then verify fields meet training requirements.
Training and test data must be kept separate. Don't mix test samples into training sets.