After completing model fine-tuning, evaluation is needed to determine whether the model has truly achieved the expected results. Since different tasks have different objectives, the evaluation methods used will also vary.
For example, classification tasks mainly focus on whether the model can determine the correct category; information extraction tasks focus on whether the required information is extracted completely and accurately; generation tasks, in addition to judging whether answers are correct, also need to focus on the completeness and relevance of content; speech recognition typically evaluates results through error rates; image generation requires a combination of automatic metrics and human evaluation.
For specific testing methods, see ใPre-evaluate Before Selecting: Using EvalScope to Form the First Report for Open-Source Modelsใ. This chapter will tell you what evaluation metrics are available for different tasks๏ฝ
Text classification is a relatively common type of task in model fine-tuning, such as intent recognition, sentiment classification, fault classification, and risk classification.
Typically, prediction results can be divided into four cases:
Accuracy represents the proportion of correctly predicted samples among all samples:
\mathrm{Accuracy}=\frac{TP+TN}{TP+TN+FP+FN}
For example, if 1000 data points are tested and the model correctly predicts 900, then the Accuracy is 90%. Accuracy can intuitively reflect the model's overall classification performance. However, when the data quantities of different categories differ significantly, relying solely on Accuracy may lead to misjudgment. If among 1000 sentiment classification data points only 20 are negative, and the model simply predicts all data as "positive," the Accuracy can still reach 98%, but in reality, none of the 20 negative data points were identified. Therefore, when class distribution is imbalanced, it is usually necessary to evaluate the model using Precision, Recall, and F1 together.
Precision represents how many of the data predicted as positive samples actually belong to positive samples:
$\mathrm{Precision}=\frac{TP}{TP+FP} $
Sign in to join the discussion
Recall represents how many of the data labeled as positive samples were successfully identified by the model:
$\mathrm{Recall}=\frac{TP}{TP+FN} $
F1 Score comprehensively considers Precision and Recall. A higher F1 typically indicates that the model has achieved a good balance between Precision and Recall.
F1 = 2 \times \frac{\mathrm{Precision} \times \mathrm{Recall}}
{\mathrm{Precision} + \mathrm{Recall}}
Accuracy, Precision, Recall, and F1 can represent the model's overall performance with a single value. However, if you want to know which categories the model tends to misclassify, you need to use a confusion matrix for further analysis. The confusion matrix compares the true categories of each test data point with the model's predicted categories, then counts the prediction results between different categories and arranges them into a matrix.
For example, if the test set contains three categories: network fault, hardware fault, and software fault, the model prediction results are as follows:
| True Category \ Predicted Category | Network Fault | Hardware Fault | Software Fault |
| Network Fault | 45 | 2 | 3 |
| Hardware Fault | 1 | 47 | 2 |
| Software Fault | 6 | 2 | 42 |
Each row represents the true category of the data, and each column represents the model's predicted category. The numbers on the matrix diagonal represent the number of correct predictions; for example, out of 50 network fault data points, 45 were correctly identified. Numbers off the diagonal represent prediction errors; for example, 6 software faults were incorrectly predicted as network faults. From the confusion matrix, you can not only see how many data points the model predicted correctly but also further discover which categories are easily confused. In the above results, software faults were relatively more frequently incorrectly predicted as network faults, indicating that the model's ability to distinguish between these two categories still has room for improvement. By analyzing these easily confused categories, you can further discover problems with the model and provide references for supplementing training data and optimizing the model.
For tasks such as named entity recognition, field extraction, and structured information extraction, commonly used evaluation metrics include field accuracy, entity-level Precision, Recall, and F1. Different metrics focus on different evaluation granularities and can be selected based on specific tasks.
Field accuracy is mainly used to measure whether the model correctly extracts specified fields. Its calculation method is:
\text{Field Accuracy} =
\frac{\text{Number of Correctly Extracted Fields}}
{\text{Total Number of Fields to Extract}}
For tasks containing multiple fields, the accuracy of different fields can be calculated separately. For example, separately calculating the accuracy of company names and person names can further analyze which fields the model has weaker extraction capabilities for. However, this method has limitations: it cannot determine which entities the model missed. Therefore, for information extraction tasks containing multiple entities, entity-level Precision, Recall, and F1 are needed for more detailed evaluation.
For Named Entity Recognition (NER) tasks, entity-level Precision, Recall, and F1 are used for evaluation. This section is similar to the classification evaluation metrics in the previous section.
Where:
Precision is used to measure how many of the entities identified by the model are correct:
Precision = \frac{TP}{TP + FP}
Recall is used to measure how many of the entities in the ground truth were successfully identified by the model:
Recall = \frac{TP}{TP + FN}
F1 comprehensively considers Precision and Recall:
F1 =
\frac{2 \times Precision \times Recall}
{Precision + Recall}
Higher Precision indicates that the entities identified by the model are generally more accurate; higher Recall indicates that the model missed fewer entities; F1 is used to comprehensively measure performance in both aspects.
Unlike classification and information extraction tasks, content generated by large language models typically has strong openness, and model outputs usually do not have a unique standard answer. This type of task usually needs to be evaluated from multiple dimensions including correctness, completeness, relevance, format compliance, and hallucination.
Correctness mainly judges whether the content generated by the model is correct, including whether there are errors in facts, conclusions, and reasoning results. For fine-tuned models, even if the response language is fluent and the format is complete, if it contains incorrect facts or conclusions, the model cannot be considered to have achieved good generation performance. Correctness is typically the most basic requirement when evaluating generation results.
Completeness mainly judges whether the model has sufficiently answered the question and whether important information that should be included has been omitted. Some responses, while not having obvious errors in content, only answer part of the question and miss key information. In such cases, the model's response has a certain degree of correctness, but completeness is still insufficient. Completeness focuses not only on whether the answer is correct but also on whether all content that should be answered has been addressed.
Relevance mainly judges whether the content generated by the model revolves around the user's question. Large language models sometimes generate a large amount of content that appears reasonable but is not closely related to the question. While this content itself may not be incorrect, it reduces the effectiveness of the response. A response with good relevance should directly address the user's question and minimize irrelevant, repetitive, or overly expanded content.
In practical applications, in addition to the response content itself, models usually need to output results in a specified format. For example, requiring the model to return results in JSON format, or organize content according to specified fields, order, and structure. Format compliance mainly evaluates whether the model generates content according to given output requirements. Even if the response content is correct, if it is not output according to the specified format, it may cause subsequent programs to be unable to parse and use it properly. For models that have undergone instruction fine-tuning or have fixed output requirements, format compliance is also an important evaluation dimension.
Hallucination refers to the model generating content that lacks basis, is inconsistent with facts, or is fabricated from nothing. When the model does not know certain information, instead of indicating insufficient information, it generates a fact that appears reasonable but does not actually exist. This is a typical hallucination problem.
Hallucination is related to correctness but focuses on different aspects. Correctness focuses on whether the response content is correct, while hallucination focuses more on whether the model has generated information without reliable basis. For scenarios such as knowledge Q&A that require high factual reliability, it is necessary to pay close attention to the model's hallucination situation.
Speech Recognition (ASR) tasks typically need to be evaluated from two aspects: recognition accuracy and processing efficiency. Among them, CER and WER are used to measure error rates in recognition results, while RTF is used to measure the speed of model processing speech.
Before introducing CER and WER, let me first explain several commonly used symbols in speech recognition error rate calculations:
Both CER and WER are calculated based on these types of errors, with the main difference being the statistical units used.
CER (Character Error Rate) measures the difference between speech recognition results and standard text using characters as units. Its calculation formula is:
CER = \frac{S + D + I}{N}
Lower CER typically indicates more accurate speech recognition results. For Chinese speech recognition tasks, CER is a commonly used evaluation metric.
WER (Word Error Rate) has essentially the same calculation method as CER:
WER = \frac{S + D + I}{N}
CER uses characters as the basic unit, while WER uses words as the basic unit. It is more commonly used in speech recognition tasks for languages like English that use words as basic linguistic units.
In addition to whether recognition results are accurate, speech recognition models in actual deployment also need to consider processing speed. RTF (Real-Time Factor) is a commonly used metric for measuring speech recognition processing efficiency. Its calculation formula is:
RTF =
\frac{\text{Model Processing Time}}
{\text{Speech Duration}}
For example, if a 10-second audio clip takes 5 seconds to recognize, then:
RTF = \frac{5}{10} = 0.5
Lower RTF indicates faster model processing speed. When RTF is less than 1, it means the model's processing speed is faster than the actual playback speed of the speech, meeting the basic conditions for real-time processing.
In addition to metrics like CER, WER, and RTF, models can also be evaluated under different testing environments. For example, calculating CER or WER separately under conditions such as background noise, different accents, distant recording, and multiple speakers. By comparing error rates under different scenarios, the model's recognition capability and stability in complex environments can be further evaluated.
Text-to-Speech (TTS) aims to convert input text into natural, clear speech. Unlike text generation models, TTS models not only need to ensure that generated content matches the input text but also need to focus on speech naturalness, timbre, and generation speed. TTS models are typically evaluated from aspects including content accuracy, speech naturalness, speaker similarity, and inference performance.
TTS first needs to ensure that input text can be correctly read aloud, avoiding misreading, missed reading, or repetition issues. During evaluation, ASR models can be used to re-recognize TTS-generated speech as text, and then compare the recognition results with the original input text. The content accuracy metrics are consistent with those in the previous ASR chapter, calculating Character Error Rate (CER) and Word Error Rate (WER).
In addition to correctly reading text, TTS-generated speech also needs to have good naturalness, such as whether the speaking rate is reasonable, whether pauses are natural, whether intonation is fluent, and whether there is a noticeable mechanical feel. Mean Opinion Score (MOS) is typically used. During evaluation, multiple testers listen to the generated speech and score it according to certain scoring standards, typically using a 1-5 scale, where 1 indicates poor speech quality and 5 indicates natural, fluent speech close to human speech.
The final MOS can be expressed as:
MOS = \frac{1}{M}\sum_{i=1}^{M}s_i
Where $M$ represents the number of people participating in the scoring, and $s_i$ represents the score given by the $i$th evaluator.
MOS can relatively intuitively reflect users' actual listening experience of generated speech, but since it relies on manual listening, its evaluation cost is relatively high and it also has a certain degree of subjectivity.
For voice cloning, multi-speaker TTS, or zero-shot TTS models, simply ensuring speech naturalness is not enough. It is also necessary to evaluate whether the generated speech maintains the target speaker's timbre characteristics. A common method is to use speaker recognition models to extract Speaker Embeddings from reference speech and generated speech separately, and then calculate the cosine similarity between the two vectors:
SIM =
\frac{
\mathbf{e}{ref} \cdot \mathbf{e}{gen}
}{
|\mathbf{e}{ref}| |\mathbf{e}{gen}|
}
Where $\mathbf{e}{ref}$ represents the speaker vector of the reference speech, and $\mathbf{e}{gen}$ represents the speaker vector of the generated speech.
Higher SIM indicates that the generated speech is closer to the reference speaker in terms of timbre characteristics. This metric is particularly suitable for evaluating the effectiveness of voice cloning models.
In terms of inference performance, TTS is similar to speech recognition models and can also use RTF to measure model processing efficiency. The difference is that RTF in speech recognition typically represents the ratio of recognition time to input speech duration, while in TTS it represents the ratio of speech generation time to generated speech duration. Lower RTF indicates higher model speech generation efficiency.
The RTF calculation formula is:
RTF =
\frac{T_{infer}}{T_{audio}}
Where $T_{infer}$ represents the inference time required to generate speech, and $T_{audio}$ represents the duration of the final generated speech.
For streaming TTS models, it is also possible to further focus on Time to First Audio, which is the time elapsed from the model receiving a text request to outputting the first playable audio segment. Lower Time to First Audio means users can hear the generated speech sooner, so this metric is particularly important for real-time speech interaction scenarios.
For tasks such as text-to-image, image-to-image, and image editing, it is difficult to comprehensively evaluate generation results using only a single metric. For example, text-to-image needs to judge whether generated images match text descriptions, image editing also needs to judge whether generated results preserve the main subjects and structure of the original image, and when evaluating the overall performance of an image generation model, it is also necessary to consider the quality and authenticity of generated images.
For text-to-image tasks, it is first necessary to judge whether generated images match the input text description. For example, if the prompt requires generating specific subjects, scenes, or styles, the images generated by the model should contain these elements as much as possible.
A common evaluation approach is to use vision-language models like CLIP to map text and images to the same feature space, and then calculate the similarity between them. CLIP was proposed by Radford et al. in 2021, training through large-scale "image-text" data to enable the model to learn the correspondence between image content and natural language. The cosine similarity between image and text can be expressed as:
Similarity =
\frac{\mathbf{v}{image}\cdot\mathbf{v}{text}}
{|\mathbf{v}{image}||\mathbf{v}{text}|}
Where
\mathbf{v}{image}
represents the image feature vector,
\mathbf{v}{text}
represents the text feature vector. Generally, higher similarity indicates that the image and text are closer in overall semantics. This type of method mainly evaluates whether overall semantics match and cannot guarantee that all details in the text are accurately generated, such as judging whether quantities, positions, and details in the image are accurate.
For image-to-image and image editing tasks, an original image or reference image is usually also provided. This type of task not only needs to judge whether generated results meet text requirements but also needs to focus on whether important content in the original image has been correctly preserved.
For example, in product image editing, the background, lighting, or overall style can be modified, but it is usually desired that important information such as the product subject, appearance structure, and logo remain consistent. At this time, the editing effectiveness can be evaluated by comparing the differences between generated images and reference images.
SSIM (Structural Similarity Index) is a classic full-reference image quality evaluation method proposed by Wang et al. in 2004. It mainly compares two images from aspects such as brightness, contrast, and structure, with particular focus on whether image structural information has changed.
In addition to SSIM, LPIPS (Learned Perceptual Image Patch Similarity) can also be used to measure perceptual differences between two images. LPIPS uses deep neural networks to extract image features, and then judges visual perceptual differences between two images based on distances between features. Related research shows that deep features can better reflect people's perception of image similarity. Although both SSIM and LPIPS require reference images, they focus on different aspects: SSIM focuses more on changes in image structure, while LPIPS focuses more on differences in visual perception. In practical tasks, appropriate metrics can be selected based on image editing objectives.
For tasks with reference images such as image editing and image-to-image, the similarity between generated images and reference images can be calculated.
The previous metrics are mainly used to compare the relationship between a generated image and text or a reference image. If you want to evaluate the overall generation performance of an image generation model, statistical analysis of a batch of generated images is usually needed.
FID (Frรฉchet Inception Distance) was proposed by Heusel et al. in 2017, mainly used to measure differences in feature distributions between generated images and real images. FID. FID first uses image models to extract features from real and generated images, then separately calculates the distributions of the two sets of features and computes the distance between the two distributions. Its basic form is:
FID =
|\mu_r-\mu_g|_2^2
+
Tr\left(
\Sigma_r+\Sigma_g
-2(\Sigma_r\Sigma_g)^{1/2}
\right)
Where $\mu_r$ and $\Sigma_r$ represent the mean and covariance matrix of real image features, and $\mu_g$ and $\Sigma_g$ represent the mean and covariance matrix of generated image features. FID compares not two specific images but the differences in feature distributions between two batches of images. Lower FID indicates that the feature distribution of generated images is closer to that of real images, making it suitable for comparing overall performance between models.
In addition to evaluating whether images match text descriptions, automatic evaluation of generated images can also be conducted from the perspective of visual effects. Aesthetic Score typically uses data containing human aesthetic ratings or preference information to train evaluation models, allowing the model to learn people's preferences for visual image quality, and then automatically score generated images.
This type of method typically comprehensively reflects characteristics such as image composition, color, lighting, clarity, and overall visual appeal. Higher scores generally indicate better visual performance of the image under the evaluation model used. Aesthetic evaluation has a certain degree of subjectivity, as different scoring models may have different training data and aesthetic preferences. Therefore, Aesthetic Score is more suitable as an auxiliary metric and cannot independently represent image generation quality.
Objective metrics can evaluate generation results from aspects such as text-image consistency, image similarity, and feature distribution, but these metrics cannot completely reflect the actual visual effects of images. Therefore, in image generation tasks, it is usually also necessary to combine human evaluation to judge the quality of generated images from the perspective of human visual perception and actual usage needs.
Human evaluation can focus on the following aspects:
To reduce differences caused by individual subjective judgments, unified scoring standards can be established in advance, and multiple evaluators can conduct evaluations separately. Subjective evaluation can discover visual problems that automated metrics have difficulty reflecting. Therefore, in actual image generation evaluation, objective metrics and subjective evaluation are usually used together to more comprehensively judge the model's generation performance.
Video generation models need to generate continuous video content based on inputs such as text and images. Compared with image generation, video not only needs to ensure the quality of individual frames but also needs to focus on continuity between different frames, reasonableness of motion, and consistency between generated content and input conditions. Therefore, video generation models usually need to be comprehensively evaluated from multiple aspects including text-video consistency, overall video quality, temporal consistency, motion quality, and subjective viewing experience.
Similar to text-to-image tasks, text-to-video tasks also need to evaluate semantic consistency between generated content and input text. The difference is that video consists of continuous video frames, requiring further consideration of the degree of matching between the entire video content and text description. The main method is to extract several video frames from generated videos, use vision-language models like CLIP to separately calculate semantic similarity between each video frame and input text, and then average the results of multiple video frames (commonly called Frame-average CLIP Score):
$\text{CLIP-Score} = \frac{1}{N} \sum_{i=1}^{N} \operatorname{Similarity}(\mathbf{v}_{\text{frame}_i}, \mathbf{v}_{\text{text}})$
Where $N$ represents the number of video frames participating in the evaluation, $\mathbf{v}_{\text{frame}_i}$ represents the feature vector of the $i$th video frame, $\mathbf{v}_{\text{text}}$ represents the feature vector of the input text, and $\operatorname{Similarity}(\cdot, \cdot)$ uses cosine similarity. Higher scores indicate higher semantic matching between video content and input text. This metric mainly focuses on the matching between spatial static content and text and cannot fully reflect action coherence and physical temporal logic in video, so it still needs to be combined with metrics such as temporal consistency and motion quality for comprehensive measurement.
Frรฉchet Video Distance (FVD) is a core metric for measuring generation distribution quality in video generation tasks, used to measure differences in deep feature distributions between generated video sets and real video sets. When calculating FVD, deep features of real and generated videos are first extracted, Gaussian distribution parameters (mean and covariance matrix) of the two sets of features are calculated, and their Frรฉchet distance is computed: $\mathrm{FVD} = \|\mu_r - \mu_g\|_2^2 + \operatorname{Tr}\left(\Sigma_r + \Sigma_g - 2\left(\Sigma_r^{1/2} \Sigma_g \Sigma_r^{1/2}\right)^{1/2}\right)$
Where $\mu_r$ and $\Sigma_r$ represent the mean vector and covariance matrix of real video features respectively, $\mu_g$ and $\Sigma_g$ represent the mean vector and covariance matrix of generated video features respectively, and $\operatorname{Tr}(\cdot)$ represents the trace of a matrix. Lower FVD values indicate that the feature distribution of generated videos is closer to real videos, indicating better overall performance of the model in visual fidelity, temporal reasonableness, and sample diversity.
Temporal Consistency is mainly used to evaluate the stability of generated videos in the time dimension. Video consists of continuous frames, requiring not only good image quality for each frame but also smooth transitions of subject appearance, background texture, color, and style between consecutive frames. For example, in character generation scenarios, if facial features or clothing textures frequently flicker or deform between frames, even if individual frame quality is high, the viewing experience will be greatly diminished. In actual evaluation, a common method is to use image feature extraction networks to obtain representations of consecutive frames and calculate the average cosine similarity between adjacent frame features. For a video containing $N$ video frames, its temporal consistency $TC$ can be expressed as:
$TC = \frac{1}{N-1} \sum_{i=1}^{N-1} \frac{\mathbf{f}_i^{\top} \mathbf{f}_{i+1}}{\|\mathbf{f}_i\|_2 \|\mathbf{f}_{i+1}\|_2}$
Where $\mathbf{f}_i$ and $\mathbf{f}_{i+1}$ represent the visual feature vectors of the $i$th and $(i+1)th frames respectively, and $N$ is the total number of sampled frames. Generally, higher $TC$ scores indicate smaller visual jumps between adjacent frames and more stable images. Temporal consistency is not absolutely better when higher: when the model degrades and generates almost completely static images, $TC$ will also be abnormally high. Temporal consistency must be combined with metrics such as Dynamic Degree and motion amplitude for joint judgment.
Video generation models not only need to avoid image flickering but also need to ensure that subject motion trajectories are coherent and natural, avoiding phenomena that do not conform to common sense such as local limb mutations and instantaneous displacement. Motion Smoothness is mainly used to measure the smoothness of changes in motion velocity and acceleration over time. In specific calculations, pre-trained optical flow estimation models (such as RAFT) are typically used to extract dense optical flow fields or motion vectors between adjacent frames. Let the motion vector (or average optical flow) from frame $t$ to frame $t+1$ be $\mathbf{v}_t$, and the motion change error $E_{\text{motion}}$ can be defined through differences in motion vectors between adjacent time periods:
$E_{\text{motion}} = \frac{1}{N-2} \sum_{t=1}^{N-2} \|\mathbf{v}_{t+1} - \mathbf{v}_t\|_2$
Where $\mathbf{v}_t$ reflects the motion state at stage $t$. Smaller $E_{\text{motion}}$ indicates fewer sudden changes in motion acceleration and smoother action transitions.
Since it is difficult for a single metric to comprehensively reflect the capabilities of video generation models, comprehensive evaluation frameworks like VBench can now be used for multi-dimensional evaluation of video generation models. VBench divides video generation quality into multiple evaluation dimensions, such as Subject Consistency, Background Consistency, Motion Smoothness, Dynamic Degree, Aesthetic Quality, and Imaging Quality. Compared with single metrics like CLIP Score and FVD, VBench's advantage lies in its ability to analyze video generation performance from multiple perspectives including image quality, temporal stability, motion performance, and condition consistency, making it more suitable for comprehensive capability comparison between different video generation models.
Large language model outputs have strong openness, and the same question can often have multiple reasonable response approaches. It is difficult to comprehensively judge whether responses are correct and meet user needs relying solely on automated metrics. Therefore, when evaluating fine-tuned large language models, human evaluation remains an important evaluation method. For open-ended content that is difficult to accurately evaluate using automated metrics, blind testing or multi-person scoring methods can be used. Evaluators can comprehensively judge model outputs from aspects such as correctness, completeness, relevance, instruction following, expression quality, and safety.
In actual evaluation, unified scoring standards can be established in advance, such as scoring different evaluation dimensions on a 1-5 scale. To reduce evaluators' preconceptions about models, model names and versions can be hidden, allowing evaluators to conduct blind testing based solely on response content. If comparing models before and after fine-tuning, responses from both models to the same question can be displayed anonymously, and evaluators can choose the better-performing response.
Since human evaluation has a certain degree of subjectivity, multi-person scoring can also be adopted for important evaluations. The same batch of content is independently evaluated by multiple evaluators, and then the scoring results are summarized. When there are large differences in scores between different evaluators, it is necessary to further check whether the evaluation standards are clear to improve the consistency and reliability of evaluation results. Although human evaluation requires more time and manpower, it can discover problems that automated metrics have difficulty reflecting. When evaluating the fine-tuning effectiveness of large language models, automated evaluation and human evaluation can be combined to more comprehensively judge whether the model has truly been improved.