AIUnlimited
๐ŸŒณ

AI Foundations

๐ŸŒฑ
AI Seeds

Start from zero

๐ŸŒฟ
AI Sprouts

Build foundations

๐ŸŒณ
AI Branches

Apply in practice

๐Ÿ•๏ธ
AI Canopy

Go deep

๐ŸŒฒ
AI Forest

Master AI

๐Ÿ”จ

AI Mastery

โœ๏ธ
AI Sketch

Start from zero

๐Ÿชจ
AI Chisel

Build foundations

โš’๏ธ
AI Craft

Apply in practice

๐Ÿ’Ž
AI Polish

Go deep

๐Ÿ†
AI Masterpiece

Master AI

๐Ÿ“˜

AI Practice

๐Ÿ“–
Understanding Open-Source Models

Fundamentals and resources for open-source models

๐ŸŽฏ
From Problem to Model Task

Converting business problems to model tasks

โšก
Running Your First Model

See your first results in 30 minutes

๐Ÿ”ง
Fine-Tuning and Evaluation

Fine-tune models and evaluate performance

๐Ÿš€
Application Systems

Build real-world AI applications

๐ŸŽจ
Generative AI

Explore open-source AIGC models

๐Ÿค–
Agents

Learn Agent frameworks and MCP tools

๐Ÿ“
Supplementary Fundamentals

LLM basics and evaluation

๐ŸŽ“

Claude Academy

๐Ÿค–
Claude 101

Learn AI basics with Claude

๐Ÿ’ป
Claude Code 101

Code with Claude as your pair programmer

๐Ÿค
Introduction to Claude Cowork

Collaborate with Claude on complex projects

โš™๏ธ
Claude Platform 101

Build apps with the Claude API

Lab

7 experiments loaded
๐ŸงฌNeural Network Sandbox๐Ÿค–AI or Human?๐Ÿฅ‹Prompt Engineering Dojo๐ŸAlgorithm Race๐Ÿง AI Trivia Challenge๐Ÿ—๏ธSystem Design Canvas
๐ŸŽฏMock InterviewEnter the Labโ†’
๐Ÿš€

Career Development

๐Ÿš€
Interview Launchpad

Start your journey

๐ŸŒŸ
Behavioral Mastery

Master soft skills

๐Ÿ’ป
Technical Interviews

Ace the coding round

๐Ÿค–
AI & ML Interviews

ML interview mastery

๐Ÿ†
Offer & Beyond

Land the best offer

Get Started
AIUnlimited

MIT Licence.

ๆฒชICPๅค‡18025655ๅท-11

Learn

  • AI Basics
  • AI Practice
  • Claude Academy
  • Lab
  • Career Development

Community

  • About
  • FAQ

Support

  • Terms of Service
  • Privacy Policy
  • Contact
AI & Engineering Academicsโ€บ๐ŸŽฏ From Problem to Model Taskโ€บLessonsโ€บModel Selection with EvalScope
๐Ÿ“‹
From Problem to Model Task โ€ข Beginnerโฑ๏ธ 20 min read

Model Selection with EvalScope

Evaluate First, Then Select: Creating Your First Report on Open-Source Models with EvalScope

This is a real and painful lesson.

Previously, when training models for business departments, I mistakenly assumed that newer models with better performance on general benchmarks would perform better in business scenarios.

When a model with outstanding performance on general benchmarks appeared, I confidently promised that I could quickly deliver a better fine-tuned model.

But that was not the case, because each model uses vastly different data in pre-training and post-training. Some models happen to have seen a lot of data similar to your business during pre-training, making fine-tuning naturally more efficient; while others, despite strong general capabilities, are simply incompatible with your domain, no matter how you fine-tune them.

Before selecting any model, general metrics can only serve as a reference. The final model selection must be based on performance on your actual business data.

Can You Pick a Good Model by Just Asking a Few Questions?

When selecting models, we often first look at model leaderboards, download counts, or online reviews. We might also directly prepare a few questions to ask the model and see how well it answers. These methods can help us quickly understand a model, but they all have certain limitations.

illustration

For example, a high download count doesn't necessarily mean the model is the most capableโ€”it could also be related to the model's release date, parameter scale, and usage barriers. Answering a few questions well also doesn't mean the model can maintain the same performance across all similar problems.

Especially when choosing between multiple models, if each model is tested with different questions, or if the criteria for judging answer quality differ, direct comparison becomes difficult. Sometimes two models both seem to answer well, but when truly tested on hundreds or even thousands of data points, there can be significant differences in accuracy and stability.

The purpose of model evaluation is to transform inconsistent evaluation standards into a more unified testing methodology. For example, prepare 500 fixed test questions, have candidate models like Qwen and DeepSeek answer them separately, then use the same rules to calculate accuracy. This way, you can compare the overall performance of different models while reducing the influence of personal judgment on the results.

Through model evaluation, we can not only see the final score of a model, but also further analyze which tasks the model performs well on, which types of questions it tends to get wrong, and how large the gap is between different models. These results can help us screen for more suitable models from multiple candidates, and also provide reference for subsequent business data testing, model fine-tuning, and deployment.

What Exactly Do Benchmark Scores Measure?

Lesson 2 of 20% complete
โ†Converting Business Problems to Model Tasks

Discussion

Sign in to join the discussion

To facilitate comparison between different models, benchmarks include various types of data, such as:

MMLU contains questions from multiple disciplines including mathematics, history, computer science, and law, primarily used to observe the model's knowledge and reasoning abilities;

illustration

C-Eval and CMMLU contain more Chinese exam and knowledge-based questions, which can be used to test the model's Chinese knowledge capabilities;

GSM8K mainly consists of math word problems, which can be used to test mathematical reasoning abilities;

illustration

HumanEval primarily tests the model's ability to write code according to requirements;

IFEval focuses on testing whether the model can complete tasks according to user-specified requirements.

illustration

In actual evaluation, you can select data that better fits your own scenario based on your business needs. For example, if the scenario focuses more on Chinese knowledge Q&A performance, you can prioritize testing datasets related to Chinese comprehension, knowledge, and reasoning; if the model is mainly used for mathematical reasoning, you can focus on testing math datasets like GSM8K.

Through these public benchmarks, you can compare the performance of different models on the same data and screen for models whose capabilities better meet your requirements. However, high benchmark scores do not necessarily mean the model will be practical in real business applications.

The data in public benchmarks are all general-purpose, while real business scenarios may contain a large amount of professional terminology and internal rules that are difficult to test through public datasets. In other words, when selecting models, you can use benchmarks for initial screening to identify models with good performance, but for scenario-specific effectiveness, you still need to use business data for evaluation.

Test Again with Your Own Business Data

General benchmarks test a model's general capabilities, while business evaluation focuses on its actual performance in specific scenarios. After the first round of screening, the next core question is clear: Is this model suitable for our business?

Business data evaluation essentially uses data generated from real business operations to validate the model. For example: for intelligent customer service, use real customer consultation records; for enterprise knowledge bases, use employees' daily retrieval and Q&A; for industrial operations, use frontline equipment fault logs.

When preparing evaluation data, people often have a misconception that more data is always better. Compared to sheer quantity, the representativeness and coverage of the data are actually more critical. If you prepare tens of thousands of data points but most are simple, repetitive questions, even high test scores won't have much reference value. A reasonable business test set should have a clear gradient: it should include daily high-frequency basic questions, as well as a certain proportion of complex long-distance questions and edge-case anomaly questions. Only through this layered real-data testing can the results be sufficiently objective.

illustration

How Can EvalScope Help with Evaluation?

EvalScope is a large model evaluation framework launched by the ModelScope community, capable of unified evaluation and horizontal comparison of different models' effectiveness and inference performance. EvalScope supports a relatively wide range, including large language models, multimodal models, and common retrieval models such as Embedding and Reranker. EvalScope's advantages are reflected in the following aspects:

  1. Balancing General Benchmarks and Custom Business Evaluation

EvalScope comes with pre-installed common public benchmark datasets such as MMLU, GSM8K, and HumanEval, allowing direct evaluation of models' general knowledge, mathematics, code, and logical reasoning capabilities. At the same time, it supports user-defined evaluation setsโ€”for example, introducing real business data such as text classification, Q&A, and information extraction into the evaluation process, thereby accurately judging model performance in specific business scenarios.

  1. Supporting Multiple Model Formats and Service Calls

The EvalScope framework supports both evaluating local offline models and standard OpenAI format interfaces. For model services deployed through mainstream inference engines like vLLM and SGLang, they can be directly integrated for evaluation, adapting to different development and deployment environments.

  1. Inference Performance Stress Testing Capability

In addition to effectiveness evaluation, the operational performance of a model after deployment is also an important consideration for selection. EvalScope can simulate multi-user concurrent requests, testing core metrics such as throughput, request latency, Time to First Token (TTFT), and Time Per Output Token (TPOT), intuitively reflecting the service's capacity and response speed under concurrent scenarios.

  1. Multi-dimensional Comparative Analysis and Report Visualization

After evaluation is complete, the EvalScope framework supports automatic generation of statistical reports and comparison charts.

Run an Evaluation to See the Model's Actual Performance

This experiment uses ModelScope's Notebook with the ubuntu22.04-cuda12.8.1-py312-torch2.10.0-1.39.0 image

illustration

Check if evalscope is already installed in the environment

!pip3 list |grep evalscope

The following output in the terminal indicates that evalscope is already installed

illustration

If not installed, you can install it directly via pip

!pip3 install evalscope
# If you need more features, use the following commands to install
!pip3 install -e '.[opencompass]'   # Install OpenCompass backend
!pip3 install -e '.[perf]'          # Install Perf dependencies
!pip3 install -e '.[app]'           # Install visualization dependencies
!pip3 install -e '.[all]'           # Install all backends (Native, OpenCompass, VLMEv)

First, Use a Public Dataset to Test Basic Capabilities

After installation, we can use a public dataset for evaluation, such as math_500. Using the Qwen3-0.6B model, for quick testing we can select only 10 data points for evaluation. The evaluation command is as follows:

!evalscope eval --model Qwen/Qwen3-0.6B --datasets math_500 --limit 10

The --model parameter specifies the model name or local model path, --datasets specifies the evaluation dataset, --limit specifies how many data points to test.

The output results are as follows:

illustration
illustration

You can view the evaluation results in the corresponding directory. eval_log.log records the evaluation results for the 10 data points.

illustration

Switch to Business Data to Test Retrieval Performance

To evaluate the model's performance on business data, we can define custom data for evaluation. This chapter uses public retrieval data from a telecommunications operator to evaluate the Qwen3-Embedding-0.6B model.

The data format is as follows:

retrieval_data1
โ”œโ”€โ”€ corpus.jsonl
โ”œโ”€โ”€ queries.jsonl
โ””โ”€โ”€ qrels.jsonl

Where corpus.jsonl is the document collection data, with the format: including {"_id": "xxx", "text": "xxx"}, where _id is the corpus ID and text is the corpus text. The format is as follows:

illustration

queries.jsonl is the query file, with each line including {"_id": "xxx", "text": "xxx"}, where _id is the query data ID and text is the query text. The format is as follows:

illustration

qrels.jsonl is the ground truth file, meaning the correct documents corresponding to each question. Each line includes {"query-id": "xxx", "corpus-id": "xxx", "score": 1}, where query-id is the query data ID, corpus-id is the document collection ID, and score is the relevance score. The format is as follows:

illustration

Simply right-click and upload these three files to the specified directory,

illustration

Install the mteb package

!pip install mteb
illustration

Install the langchain_core package

!pip install langchain_core
illustration

Next, we can evaluate the Embedding retrieval performance. The code is as follows:

from evalscope.run import run_task

task_cfg = {
    "work_dir": "outputs",

    "eval_backend": "RAGEval",

    "eval_config": {

        "tool": "MTEB",

        "models": [
            {
                "model_name_or_path": "Qwen/Qwen3-Embedding-0.6B",

                "max_seq_length": 512,

                "model_kwargs": {
                    "torch_dtype": "auto"
                },

                "encode_kwargs": {
                    "batch_size": 8,
                },
            }
        ],

        "eval": {

            "custom_tasks": [
                {
                    "name": "MyRetrieval",
                    "data_path": "./retrieval_data1"
                }
            ],

            "overwrite_results": True,

            "verbosity": 2
        },
    },
}

run_task(task_cfg=task_cfg)

The execution results are as follows:

illustration

Not Just Effectiveness Testingโ€”Performance Testing Too

After completing the model effectiveness test, you can continue with performance stress testing. If the model is already deployed through inference frameworks like vLLM, you can use EvalScope to send requests to the model service with different concurrency levels.

For example, you can set different concurrency levels of 1, 10, 50, and 100 to observe the model's throughput, first token latency, and request success rate as concurrency increases. EvalScope's performance test results can track metrics such as RPS, output throughput, average latency, and P50/P99, and save the results as a test report.

Under this image, the default package installation may cause errors. You can resolve this by downgrading the modelscope version:

!python -m pip install --no-cache-dir --force-reinstall "modelscope==1.37.1"

Open the Notebook and start Qwen3.5-2B with vLLM in the terminal:

CUDA_VISIBLE_DEVICES=0 \
vllm serve Qwen/Qwen3-0.6B \
    --served-model-name qwen3-0.6b \
    --host 0.0.0.0 \
    --port 8000 \
    --gpu-memory-utilization 0.6
illustration
illustration

After the model starts, you need to test if the API is accessible. Enter the following code:

import requests

url = "http://127.0.0.1:8000/v1/chat/completions"

data = {
    "model": "qwen3-0.6b",
    "messages": [
        {
            "role": "user",
            "content": "Hello, please introduce yourself"
        }
    ],
    "max_tokens": 100
}

response = requests.post(url, json=data)

print("Status code:", response.status_code)
print(response.text)

If the API is accessible, it will display the model output:

illustration

Next, use EvalScope to evaluate performance. The metrics we typically focus on include Time to First Token (TTFT) and Time Per Output Token (TPOT).

!evalscope perf \
    --model qwen3-0.6b \
    --url http://127.0.0.1:8000/v1/chat/completions \
    --api openai \
    --api-key sk-123456 \
    --dataset random \
    --tokenizer-path Qwen/Qwen3-0.6B \
    --min-prompt-length 1024 \
    --max-prompt-length 1024 \
    --min-tokens 1024 \
    --max-tokens 1024 \
    --parallel 1 2 4 8 \
    --number 10 20 50 100

Here are the parameter descriptions:

  • --model qwen3.5-2b: Model name, must match the model name configured in the vLLM inference service
  • --url http://127.0.0.1:8000/v1/chat/completions: The actual API endpoint of the model service
  • --api-key sk-123456: API authentication key. Must be provided if the server requires authentication; if not configured, this parameter can be omitted.
  • --dataset random: Specifies the test dataset in random generation mode, no need to prepare local test files.
  • --tokenizer-path Qwen/Qwen3.5-2B: Specifies the path or Hugging Face model name of the Tokenizer, used to precisely control the Token length of generated text.
  • --min-prompt-length 1024 and --max-prompt-length 1024: Control the length range of input prompts. Both set to 1024, meaning each request has a fixed input length of 1024 Tokens.
  • --min-tokens 1024 and --max-tokens 1024: Control the output length range generated by the model. Both set to 1024, forcing the model to output a fixed 1024 Tokens each time (ignoring end tokens).
  • --parallel 1 2 4 8: Set concurrency level gradients. During stress testing, it will sequentially simulate 1, 2, 4, and 8 users making requests simultaneously to observe system performance under different concurrency pressures.
  • --number 10 20 50 100: Set the total number of requests for each concurrency level. This list corresponds one-to-one with --parallel. For example, at concurrency 1, a total of 10 requests are sent; at concurrency 8, a total of 100 requests are sent.
illustration

The benchmark.log results are as follows:

illustration

EvalScope provides us with a method for quickly evaluating models. You can view the relevant HTML report content in the corresponding output folder. For more EvalScope usage, visit https://evalscope.readthedocs.io/zh-cn/v1.6.1/get_started/introduction.html.

For the specific experiment process, refer to: https://modelscope.cn/gallery/liucong/5a794182-147f-4c97-83c5-233ca6965f4b

Bonus

Try using your own real business data to test the effectiveness of current open-source models!