على الرغم من أن النماذج الكبيرة العامة تمتلك قدرات قوية على أسئلة وأجوبة المعرفة، إلا أن التطبيقات العملية في المؤسسات تتضمن العديد من الأسئلة المتعلقة بالمواد الداخلية للمنتجات، والوثائق التقنية، والسياسات التشغيلية، ودليل الإجراءات. عادةً ما تكون هذه المعرفة غير مدرجة في بيانات تدريب النموذج، وهي قابلة للتحديث المستمر مع تطور الأعمال.
توفر RAG (التوليد المعزز بالاسترجاع) حلاً يقوم باسترجاع المحتوى ذي الصلة بالسؤال من قاعدة معرفة المؤسسة قبل أن يجيب النموذج الكبير، ثم تقديم نتائج الاسترجاع إلى النموذج الكبير لتوليد الإجابة.
يمكن تلخيص الفكرة الأساسية لـ RAG كالتالي: استرجاع أولاً، ثم الإجابة. تعتمد أسئلة وأجوبة النماذج الكبيرة التقليدية بشكل أساسي على المعرفة المدربة بالفعل في معلمات النموذج. على سبيل المثال، عندما يطرح المستخدم سؤالاً، يقوم النموذج بإنشاء إجابة مباشرة بناءً على المعلومات المكتسبة أثناء التدريب. تضيف RAG خطوة استرجاع المعرفة إلى هذه العملية. عندما يطرح المستخدم سؤالاً، يقوم النظام أولاً بالبحث في قاعدة المعرفة الخارجية عن المحتوى ذي الصلة بالسؤال، ثم يرسل سؤال المستخدم والمعرفة المسترجعة معاً إلى النموذج الكبير، الذي يقوم بتوليد الإجابة النهائية عن طريق دمج هذا المحتوى.
يمكن تمثيل سير عمل RAG الأساسي كالتالي: سؤال المستخدم ← استرجاع المعرفة ← استرجاع المحتوى ذي الصلة ← نموذج LLM يولّد الإجابة ← إعادة النتيجة. لا تجعل RAG النموذج يعيد تعلم هذه المعرفة؛ بدلاً من ذلك، توفر مؤقتاً للمواد المرجعية اللازمة للسؤال الحالي عند الإجابة.
ليست RAG مطلوبة لجميع تطبيقات النماذج الكبيرة، لكنها شائعة جداً في سيناريوهات مثل أسئلة وأجوبة المعرفة المؤسسية، وخدمة العملاء الذكية، ومساعدات المستندات.
إذا كانت المعرفة التجارية تتطلب تحديثات مستمرة، فإن RAG عادةً ما تكون مناسبة. على سبيل المثال، قد يتم تعديل وصف المنتجات والسياسات التجارية وإرشادات التشغيل للشركة باستمرار. إذا تم كتابة هذه المعرفة في النموذج من خلال تدريب النموذج، فإن إعادة تدريب النموذج في كل مرة تتغير فيها المعرفة تكون مكلفة وصعبة الصيانة.
تسمح RAG بالتحديثات المباشرة لقاعدة المعرفة الخارجية. عند تغيير المستندات، تحتاج فقط إلى إعادة معالجة المستندات ذات الصلة وتحديث الفهرس، دون إعادة تدريب النموذج الكبير بالكامل. إذا كانت الإجابات تحتاج إلى تقديم مراجع المصدر، فإن RAG مناسبة أيضاً. نظراً لأن إجابات النموذج مولّدة بناءً على المستندات المسترجعة، يمكنها إسماع أسماء المستندات والأقسام وأرقام الصفحات أو مقتطفات النص الأصلي في وقت واحد، مما يسمح للمستخدمين بمعرفة مصدر الإجابة.
يمكن استخدام كل من RAG والضبط الدقيق للنموذج لتحسين أداء النماذج الكبيرة في الأعمال العملية، لكن النهجين يحلان المشكلات بشكل مختلف. تعالج RAG بشكل أساسي مشكلة "نقص النموذج في المعرفة ذات الصلة"، بينما يعالج الضبط الدقيق للنموذج بشكل أساسي مشكلة "عدم قدرة النموذج على إكمال المهام وفقاً للمتطلبات".
لا تُعدّل RAG معلمات النموذج نفسها. بدلاً من ذلك، قبل أن يجيب النموذج على سؤال، يقوم أولاً باسترجاع المحتوى ذي الصلة بالسؤال من قاعدة المعرفة الخارجية، ثم يقدم هذا المحتوى إلى النموذج كأساس للإجابة.
تسجيل الدخول للانضمام إلى النقاش
يعمل الضبط الدقيق للنموذج بشكل مختلف. يتطلب الضبط الدقيق استخدام بيانات تدريب محددة لتدريب النموذج بشكل أقصى، وخلال التدريب، يتم تعديل معلمات النموذج حتى يتعلم النموذج تدريجياً كيفية التعامل مع مهام معينة. على سبيل المثال، تتطلب مهمة التعرف على الكيانات المذكورة في الفصل 9 أن يقوم النموذج بإخراج إجابات بتنسيق ثابت. عندما تتغير متطلبات المهام بشكل كبير، عادةً ما يكون من الضروري إعداد بيانات تدريب جديدة وإعادة الضبط الدقيق.
في التطبيق العملي، يمكنك اختيار طرق مختلفة بناءً على المشكلة المراد حلها: عندما تكون هناك حاجة إلى تكملة أو تحديث المعرفة، فإن RAG أكثر ملاءمة؛ عندما تكون هناك حاجة إلى تعديل قدرات المهمة للنموذج أو تنسيق الإجابة أو سلوكه، فإن الضبط الدقيق للنموذج أكثر ملاءمة. يمكن أيضاً استخدام RAG والضبط الدقيق للنموذج معاً. على سبيل المثال، يمكن للضبط الدقيق أن يساعد النموذج على إتقان قدرات الأسئلة والأجوبة أو اتباع التعليمات، ثم يمكن لـ RAG توفير المعرفة الخاصة بالمجال، مما يمكّن النموذج من إكمال المهام وفقاً لمتطلبات الأعمال وتوليد إجابات بناءً على أحدث معرفة في المجال.
A complete RAG system is generally divided into two main phases: knowledge base construction and online Q&A. Knowledge base construction is mainly responsible for processing raw materials such as PDFs, Word documents, Markdown files, and web pages into searchable knowledge. Online Q&A involves finding relevant content from the established knowledge base after the user asks a question. This experiment uses the publicly available "Regulations on the Implementation of the Road Traffic Safety Law of the People's Republic of China" as the document source to build a RAG system for road traffic safety law Q&A.
عادةً ما تكون معرفة المؤسسة موزعة عبر أنواع ملفات مختلفة مثل PDF وWord وExcel وMarkdown وصفحات الويب. نظراً للاختلافات الكبيرة في التنسيق والبنية الداخلية بين الملفات المختلفة، قبل بناء قاعدة معرفة RAG، عادةً ما تكون هذه المستندات بحاجة إلى تحويل إلى نص موحد أو بيانات منظمة قابلة للمعالجة. تسمى هذه العملية بتحليل المستند. For RAG systems, a good document parsing tool not only recognizes text characters but also identifies and preserves structural information from the original document, such as: document titles, paragraphs, tables, images, formulas, and more.
There are already many open-source tools available for document parsing, with the most common ones including:
1) MinerU is an open-source parsing tool designed for complex documents, capable of converting PDFs, images, Word documents, PPTs, Excel files, and other documents into machine-readable formats such as Markdown and JSON. It can recognize titles, body text, tables, images, formulas, and other content, and restore the document structure as much as possible following human reading order, making it well-suited as a document preprocessing tool in RAG systems, opendatalab/MinerU
2) MonkeyOCR is a document parsing project based on multimodal models that parses documents through structure recognition, content recognition, and relationship modeling. It can handle complex content such as text, tables, and formulas, and supports both Chinese and English documents. For PDFs with complex layouts where traditional text extraction tools perform poorly, this type of vision-model-based parsing approach can be considered, Yuliang-Liu/MonkeyOCR
3) Dolphin is a document image parsing model open-sourced by ByteDance that processes documents using an "analyze first, parse second" approach. It first identifies page layout and reading order, then further parses different types of document elements such as text, tables, formulas, and code. It is suitable for processing PDFs with complex layouts or scanned documents, ByteDance/Dolphin
4) PaddleOCR is an OCR and document parsing tool open-sourced by PaddlePaddle. In addition to common text recognition, it also provides layout analysis, table recognition, and structured document parsing capabilities, converting content from PDFs and images into structured data more suitable for downstream AI system processing. For scanned PDFs, image-based documents, and documents containing large amounts of Chinese text, PaddleOCR is a common choice, PaddlePaddle/PaddleOCR
Different document parsing tools each have their own characteristics, and no single tool is suitable for all documents. When actually building a RAG system, you can select an appropriate parsing tool based on factors such as document type, layout complexity, parsing accuracy, and deployment cost.
This experiment uses MinerU for document parsing. First, install MinerU in the ModelScope Notebook environment by running the following command:
pip install -U "mineru[all]"

After installation, you can directly parse documents by entering the file path to be parsed. Run the following command:
mineru -p /mnt/workspace/RAG/data -o /mnt/workspace/RAG/output
Where: -p is the input file path, and -o is the model output file path, which is the parsed text.

The parsed files are shown below, including many intermediate files that can be selected as needed. In this experiment, we choose the md file as the document source, which includes paragraph format information of the text.

بعد إكمال تحليل المستند، تحتاج المستندات الأطول إلى تقسيمها إلى أجزاء نصية أصغر، تُعرف بشكل شائع باسم Chunks. In RAG systems, Embedding and retrieval typically use Chunks as the basic unit, so proper document splitting helps improve the accuracy of subsequent retrieval.
For parsed Markdown documents, you can first split the content into paragraphs based on blank lines, then merge adjacent paragraphs sequentially according to the set chunk_size. When the content exceeds the specified length, a new Chunk is created.
To avoid context information loss at split points, you can also set chunk_overlap to retain a small amount of duplicate content between adjacent Chunks. For example: chunk_size = 500, chunk_overlap = 50 means each Chunk is controlled to approximately 500 characters, with adjacent Chunks retaining about 50 characters of overlapping content. After splitting, each Chunk can save information such as id, content, and file for subsequent vectorization, retrieval, and answer source localization.
فيما يلي شفرة تقسيم Chunks:
import os
import json
def load_markdown(file_path):
with open(file_path, "r", encoding="utf-8") as f:
return f.read()
def split_markdown(text, chunk_size=500, chunk_overlap=50):
"""
Split Markdown text into multiple Chunks
chunk_size: approximately how many characters each Chunk contains
chunk_overlap: how many duplicate characters to retain between adjacent Chunks
"""
paragraphs = text.split("\n\n")
chunks = []
current_chunk = ""
for paragraph in paragraphs:
paragraph = paragraph.strip()
if not paragraph:
continue
if len(current_chunk) + len(paragraph) <= chunk_size:
if current_chunk:
current_chunk += "\n\n" + paragraph
else:
current_chunk = paragraph
else:
if current_chunk:
chunks.append(current_chunk)
overlap_text = current_chunk[-chunk_overlap:] if current_chunk else ""
current_chunk = overlap_text + "\n\n" + paragraph
if current_chunk:
chunks.append(current_chunk)
return chunks
def save_chunks(chunks, source_file, output_file):
"""
Save Chunks as JSONL file
"""
file_name = os.path.basename(source_file)
with open(output_file, "w", encoding="utf-8") as f:
for i, chunk in enumerate(chunks):
data = {"id": i,"content": chunk,"file": file_name
}
f.write(
json.dumps(data, ensure_ascii=False) + "\n"
)
if __name__ == "__main__":
file_path = "./output/road-safety-implementation-regulations/hybrid_auto/road-safety-implementation-regulations.md"
text = load_markdown(file_path)
chunks = split_markdown(
text,
chunk_size=500,
chunk_overlap=50
)
print("Chunk count:", len(chunks)
save_chunks(
chunks,
source_file=file_path,
output_file="output/chunks.jsonl"
)
After execution, you can see that the previously parsed md file has been split into 39 chunks.

After completing document splitting, you need to further build a searchable document library. Since computers cannot directly search based on natural language semantics, an Embedding model is needed to convert each Chunk into a vector representation. Embedding models can map text to a high-dimensional vector space, where semantically similar text has vectors that are closer together in the space. Therefore, after a user asks a question, the question can also be converted into a vector, which is then compared with the Chunk vectors in the document library to retrieve the content most semantically relevant to the question.
The open-source community has already provided various Embedding models, such as BGE-M3 and Qwen3-Embedding. BGE-M3 is a multilingual Embedding model released by BAAI, supporting over 100 languages and text input up to 8192 Tokens. Compared to ordinary Dense Embedding models, BGE-M3 simultaneously supports dense retrieval, sparse retrieval, and Multi-Vector retrieval, making it suitable for both common vector semantic retrieval and hybrid retrieval scenarios FlagEmbedding.
Qwen3-Embedding is a text vector model series released by the Qwen team, built on the Qwen3 architecture, mainly targeting tasks such as text retrieval, text clustering, text classification, and code retrieval. Qwen3-Embedding provides different parameter scales including 0.6B, 4B, and 8B, which can be selected based on model performance and computational resources, while having good multilingual and long text processing capabilities Qwen3-Embedding.
This section uses Qwen3-Embedding-0.6B as the Embedding model for experimentation. Chapter 9 already introduced ms-swift; in addition to large model training and fine-tuning, ms-swift also supports training and inference for the Qwen3-Embedding series. Therefore, this section continues to use ms-swift to launch Qwen3-Embedding-0.6B, and completes vectorization of document Chunks and user questions through the API.
يمكنك تشغيل خدمة Embedding باستخدام الأمر التالي:
CUDA_VISIBLE_DEVICES=0 \
swift deploy \
--model Qwen/Qwen3-Embedding-0.6B \
--task_type embedding \
--vllm_gpu_memory_utilization 0.2 \
--vllm_max_model_len 512 \
--host 0.0.0.0 \
--port 18000
بعد تشغيل الخدمة، تظهر المعلومات التالية:

Commonly used vector retrieval tools include Faiss and Milvus. Faiss is an open-source vector similarity retrieval library from Meta, mainly used for efficient vector indexing and approximate neighbor search. It is simple to deploy, does not require starting a separate database service, and can be used directly in Python programs, making it well-suited for learning, experimentation, and small-to-medium-scale RAG systems. Milvus is a vector database designed for large-scale vector data. In addition to vector retrieval, it also provides data persistence, distributed storage, index management, and various query capabilities, making it more suitable for scenarios with large data scales, long-term operation, and production-level deployment. Compared to Faiss, Milvus has more complete functionality but a relatively more complex deployment and usage process.
يستخدم هذا التجربة Faiss لتخزين المتجهات واسترجاع التشابه. First, install the faiss library by running the following command:
pip install faiss-cpu

After installation, you can send the Chunks generated in the previous section to the Embedding API one by one to obtain the corresponding vectors, and establish a mapping between vectors and Chunk metadata such as id, content, and file. This enables fast similarity calculation and knowledge retrieval later.
فيما يلي شفرة بناء الفهرس build_index.py:
import json
import requests
import numpy as np
import faiss
EMBEDDING_URL = "http://localhost:18000/v1/embeddings"
MODEL_NAME = "Qwen3-Embedding-0.6B"
def get_embedding(text):
payload = {"model": MODEL_NAME,"input": text}
response = requests.post(
EMBEDDING_URL,
json=payload,
timeout=60
)
response.raise_for_status()
data = response.json()
embedding = data["data"][0]["embedding"]
return np.array(embedding, dtype="float32")
def load_chunks(file_path):
chunks = []
with open(file_path, "r", encoding="utf-8") as f:
for line in f:
chunks.append(json.loads(line)
return chunks
def build_faiss_index(chunks,index_path="faiss.index", metadata_path="metadata.json"):
embeddings = []
for i, chunk in enumerate(chunks):
text = chunk["content"]
embedding = get_embedding(text)
embeddings.append(embedding)
print(f"Processed {i + 1}/{len(chunks)}")
embeddings = np.array(embeddings, dtype="float32")
faiss.normalize_L2(embeddings)
dimension = embeddings.shape[1]
index = faiss.IndexFlatIP(dimension)
index.add(embeddings)
print("Vector count:", index.ntotal)
print("Vector dimension:", dimension)
faiss.write_index(index, index_path)
with open(metadata_path, "w", encoding="utf-8") as f:
json.dump(chunks,f,ensure_ascii=False,indent=2
)
print(f"Faiss index saved to: {index_path}")
print(f"Chunk metadata saved to: {metadata_path}")
if __name__ == "__main__":
chunks = load_chunks("output/chunks.jsonl")
build_faiss_index(
chunks,
index_path="output/faiss.index",
metadata_path="output/metadata.json"
)
بعد التنفيذ، يتم إخراج ملفات الفهرس والمتجهات المقابلة.

Embedding models can convert user questions and document Chunks into vectors respectively, and perform retrieval based on the distance or similarity between vectors. The advantage of this approach is that document vectors can be pre-computed and stored, and when a user asks a question, only one question vector needs to be computed to quickly recall relevant content from a large number of documents. However, this efficient retrieval method also has certain limitations. Embedding models encode Queries and Chunks separately, and when computing relevance, they compare the similarity of two vectors in the vector space. Reranker models use interactive matching. By inputting both the Query and candidate Chunk into the model simultaneously, the model directly analyzes the relevance between the two text segments, enabling more detailed semantic matching.
Therefore, in RAG systems, a two-stage retrieval approach of Embedding recall + Reranker re-ranking is typically used: first, Embedding is used to quickly screen a batch of candidate results from a large number of Chunks, then Reranker is used to perform more precise relevance judgment and re-ranking on a small number of candidate results. This ensures retrieval efficiency while further improving the quality of context ultimately provided to the large language model.
يستخدم هذا القسم Qwen3-Reranker-0.6B كنموذج لإعادة الترتيب، ويمكنك استخدام ms-swift لتشغيل الخدمة:
CUDA_VISIBLE_DEVICES=0 \
swift deploy \
--model Qwen/Qwen3-Reranker-0.6B \
--task_type generative_reranker \
--infer_backend transformers \
--host 0.0.0.0 \
--port 18001

بعد تشغيل الخدمة، يمكنك إدخال سؤال المستخدم وChunks المرشحة المسترجعة بواسطة Embedding إلى Reranker، وإعادة ترتيبها بناءً على درجات الصلة المحسوبة بواسطة النموذج، واختيار Chunks الأعلى تصنيفاً كمحتوى مرجعي لتوليد إجابة نموذج اللغة الكبير اللاحق.
The Embedding service, Reranker service, and Faiss vector index construction have been completed above. Based on this, candidate Chunks are screened using the two-stage retrieval approach of Embedding recall + Reranker re-ranking.
في المرحلة الأولى، يُستخدم Faiss لاسترجاع المتجهات من قاعدة المعرفة. First, the top 10 candidate Chunks by similarity ranking are obtained. Below are the top 10 vector retrieval results for the question: "For initial motor vehicle license plate and driving license application, which department should be approached for registration?"
import json
import faiss
from build_index import get_embedding
def load_faiss_index(index_path="faiss.index",
metadata_path="metadata.json"):
index = faiss.read_index(index_path)
with open(metadata_path, "r", encoding="utf-8") as f:
metadata = json.load(f)
return index, metadata
def search(query,index,metadata,top_k=10):
query_embedding = get_embedding(query)
query_embedding = query_embedding.reshape(1, -1)
faiss.normalize_L2(query_embedding)
scores, indices = index.search(query_embedding,top_k)
results = []
for score, idx in zip(scores[0], indices[0]):
if idx == -1:
continue
chunk = metadata[idx]
results.append({
"score": float(score),
"id": chunk.get("id"),
"content": chunk.get("content"),
"file": chunk.get("file")
})
return results
index, metadata = load_faiss_index(
"output/faiss.index",
"output/metadata.json"
)
query="For initial motor vehicle license plate and driving license application, which department should be approached for registration?"
candidates = search(
query=query,
index=index,
metadata=metadata,
top_k=10
)
print("candidates",candidates)

في المرحلة الثانية، يُستخدم Reranker لإعادة ترتيب النتائج المرشحة. The user question and recalled candidate Chunks are submitted to the Reranker, which recalculates the relevance scores between them and sorts them from highest to lowest.
def rerank(query, candidates):
results = []
for candidate in candidates:
response = rerank_client.chat.completions.create(
model="Qwen3-Reranker-0.6B",
messages=[
{
"role": "user",
"content": query
},
{
"role": "assistant",
"content": candidate["content"]
}
]
)
score = response.choices[0].message.content[0]
item = candidate.copy()
item["rerank_score"] = float(score)
results.append(item)
results.sort(
key=lambda x: x["rerank_score"],
reverse=True
)
for rank, item in enumerate(results, start=1):
item["rerank_rank"] = rank
return results
Below are the results after re-ranking by the Reranker model for the question: "For initial motor vehicle license plate and driving license application, which department should be approached for registration?"

From the two results above, it can be seen that after re-ranking, the Chunk order has changed significantly. In practical projects, adjustments can be made based on knowledge base size, Chunk length, question complexity, and actual retrieval performance. For example, for questions whose answers are distributed across multiple document fragments, you canappropriately increase the number of Chunks retained in the final results; you can also set a relevance threshold based on Reranker scores to further filter out low-relevance content. In addition to the above approach, practical RAG systems can also use hybrid retrieval (Hybrid Search), combining vector retrieval with keyword retrieval methods like BM25, fusing results from multiple recall routes before re-ranking.
After completing retrieval and re-ranking, several Chunks with high relevance to the user's question have been obtained. Next, these contents need to be submitted to the LLM along with the user's question as reference material. If relevance meets the requirements, the retrieved content is submitted to the large language model to generate an answer; if no sufficiently relevant content is found in the knowledge base, a refusal response is returned directly, avoiding the model generating answers without sufficient basis.
أولاً، قم بتشغيل خدمة النموذج الكبير في الـ Notebook. This experiment uses the Qwen3-4B model, started as follows:
CUDA_VISIBLE_DEVICES=0 swift deploy \
--model Qwen/Qwen3-4B \
--load_args false \
--infer_backend vllm \
--enable_thinking false \
--host 0.0.0.0 \
--port 18002 \
--api_key 123 \
--vllm_gpu_memory_utilization 0.8 \
--vllm_max_model_len 8000 \
--max_new_tokens 2000
بعد التشغيل، تظهر المعلومات التالية، مما يشير إلى بدء تشغيل النموذج بنجاح. After the service starts, you can call Qwen3-4B through port 18002.

يعيد Reranker الترتيب بناءً على صلة سؤال المستخدم مع Chunks المرشحة. This experiment directly selects the top 5 ranked Chunks as reference material for the LLM. To make the model answer as much as possible based on the knowledge base content, the answer scope can be clearly defined in the Prompt. Meanwhile, when the provided reference material cannot answer the user's question, the model is required to directly refuse to answer rather than supplementing the answer with its own knowledge. The prompt is as follows:
prompt = f"""
Please answer the user's question based on the reference material below.
Reference material:
{context}
User question:
{query}
Requirements:
1. Answer the question only based on the reference material;
2. Answer accurately and concisely;
3. Do not supplement information not present in the reference material;
4. If the reference material cannot answer this question, please answer:
"This question cannot be answered with the current knowledge base."
"""
بعد توليد النموذج الكبير للإجابة، بالإضافة إلى إعادة الإجابة النهائية إلى المستخدم، يمكن أيضاً إعادة مصادر المستند المرجعية لهذه الإجابة. For enterprise knowledge Q&A scenarios, source information helps users understand which knowledge base documents the answer comes from, and when further confirmation is needed, they can return to the original document to view the relevant content. Source information does not need to be generated by the large model, but is directly obtained from the metadata of the retrieval results. When constructing the knowledge base earlier, each Chunk retained its corresponding file field, so file names can be extracted from the top 5 Reranker-ranked Chunks, and duplicate files can be deduplicated.
Below is the core code for calling the large model to implement Q&A:
llm_client = OpenAI(
api_key="123",
base_url="http://127.0.0.1:18002/v1"
)
LLM_MODEL = "Qwen3-4B"
def generate_answer(query, top_chunks):
context = "\n\n".join(
[
f"[Reference{i + 1}]\n{item['content']}"
for i, item in enumerate(top_chunks)
]
)
prompt = f"""
Please answer the user's question based on the reference material provided below.
Reference material:
{context}
User question:
{query}
Requirements:
1. Answer the question only based on the content in the reference material;
2. The answer should be accurate and concise, do not supplement information not present in the reference material;
3. If there is no answer related to the question in the reference material, please answer:
"This question cannot be answered with the current knowledge base."
"""
response = llm_client.chat.completions.create(
model=LLM_MODEL,
messages=[{"role": "user","content": prompt }],
temperature=0.1,
max_tokens=512
)
return response.choices[0].message.content
في هذه المرحلة، تم بناء مساعد أسئلة وأجوبة مؤسسي بنجاح. أدناه يمكننا إجراء اختبار بسيط في الطرفية، ونتائج الاختبار كالتالي:

بعد إكمال نظام RAG، لا يزال يلزم تقييم لتحديد ما إذا كان عملية الأسئلة والأجوبة الكاملة فعالة فعلياً. Unlike ordinary large model Q&A, RAG results are affected by both the knowledge retrieval and answer generation stages. Therefore, RAG evaluation usually needs to be conducted from both aspects separately.
The evaluation criterion for the retrieval stage is whether the user's question can find an answer in the recalled Chunks. The commonly used metric is Recall@K. Recall@K is a relatively intuitive metric used to measure how many relevant documents the top K retrieval results cover:
Recall@K: Whether the correct document appears in the top K retrieval results.
Recall@K =
\frac{\text{Number of relevant documents retrieved in Top K}}
{\text{Total number of relevant documents}}
For example, if a question corresponds to 2 correct knowledge fragments, and both are found in the top 5 retrieval results, then Recall@5 is 100%.
Retrieving relevant knowledge does not necessarily mean the final answer is correct. It is also necessary to further evaluate the LLM's generation results. The generation stage can focus on answer accuracy, completeness, relevance, citation accuracy, and hallucination. This part can refer to Section 12.3 "Generation Task Evaluation Metrics" in Chapter 12. Through these metrics, you can determine whether the model can accurately use retrieved knowledge to generate answers, and further discover issues such as answer omissions, content deviation, or unsupported generation. In practical applications, if retrieval results are already quite accurate but generation performance still cannot meet requirements, you can consider choosing a model with stronger capabilities and larger parameter scale as the RAG generation model to improve the understanding and answering ability for complex questions.
يمكن العثور على الكود والملفات ذات الصلة بهذا الفصل على: https://modelscope.cn/gallery/liucong/a895ace8-420c-4421-ba43-4e3194392a95