OliveSensorAPI/evaluate/General_evaluation_EN.md

# EmoLLM's general evaluation

## Introduction

This document provides instructions on how to use the 'eval.py' and 'metric.py' scripts. These scripts are used to evaluate the generation results of EmoLLM- a large model of mental health.

## Installation

- Python 3.x
- PyTorch
- Transformers
- Datasets
- NLTK
- Rouge
- Jieba

It can be installed using the following command:

```bash
pip install torch transformers datasets nltk rouge jieba
```

## Usage

### convert.py

Convert raw multi-round conversation data into single round data for evaluation.

### eval.py

The `eval.py` script is used to generate the doctor's response and evaluate it, mainly divided into the following parts:

1. Load the model and word divider.
2. Set test parameters, such as the number of test data and batch size.
3. Obtain data.
4. Generate responses and evaluate.

### metric.py

The `metric.py` script contains functions to calculate evaluation metrics, which can be set to evaluate by character level or word level, currently including BLEU and ROUGE scores.

## Results

Test the data in data.json with the following results:

| Model    | ROUGE-1 | ROUGE-2 | ROUGE-L | BLEU-1  | BLEU-2  | BLEU-3  | BLEU-4  |
|----------|---------|---------|---------|---------|---------|---------|---------|
| Qwen1_5-0_5B-chat | 27.23%  | 8.55%   | 17.05%  | 26.65%  | 13.11%  | 7.19%   | 4.05%   |
| InternLM2_7B_chat_qlora | 37.86%  | 15.23%   | 24.34%  | 39.71%  | 22.66%  | 14.26%   | 9.21%   |
| InternLM2_7B_chat_full  | 32.45%  | 10.82%   | 20.17%  | 30.48%  | 15.67%  | 8.84%   | 5.02%   |
| InternLM2_7B_base_qlora_5epoch  | 41.94%  | 20.21%   | 29.67%  | 42.98%  | 27.07%  | 19.33%   | 14.62%   |
| InternLM2_7B_base_qlora_10epoch | 43.47%  | 22.06%   | 31.4%  | 44.81%  | 29.15%  | 21.44%   | 16.72%   |
Update General_evaluation_EN.md 2024-03-11 15:37:50 +08:00			`# EmoLLM's general evaluation`
新增ENmd文档 2024-03-10 15:52:18 +08:00
			`## Introduction`

			`This document provides instructions on how to use the 'eval.py' and 'metric.py' scripts. These scripts are used to evaluate the generation results of EmoLLM- a large model of mental health.`

			`## Installation`

			`- Python 3.x`
			`- PyTorch`
			`- Transformers`
			`- Datasets`
			`- NLTK`
			`- Rouge`
			`- Jieba`

			`It can be installed using the following command:`

			```bash
			`pip install torch transformers datasets nltk rouge jieba`
			```

			`## Usage`

			`### convert.py`

			`Convert raw multi-round conversation data into single round data for evaluation.`

			`### eval.py`

			The `eval.py` script is used to generate the doctor's response and evaluate it, mainly divided into the following parts:

			`1. Load the model and word divider.`
			`2. Set test parameters, such as the number of test data and batch size.`
			`3. Obtain data.`
			`4. Generate responses and evaluate.`

			`### metric.py`

			The `metric.py` script contains functions to calculate evaluation metrics, which can be set to evaluate by character level or word level, currently including BLEU and ROUGE scores.

Update General_evaluation_EN.md 2024-03-11 15:28:56 +08:00			`## Results`
新增ENmd文档 2024-03-10 15:52:18 +08:00
			`Test the data in data.json with the following results:`

			`\| Model \| ROUGE-1 \| ROUGE-2 \| ROUGE-L \| BLEU-1 \| BLEU-2 \| BLEU-3 \| BLEU-4 \|`
			`\|----------\|---------\|---------\|---------\|---------\|---------\|---------\|---------\|`
			`\| Qwen1_5-0_5B-chat \| 27.23% \| 8.55% \| 17.05% \| 26.65% \| 13.11% \| 7.19% \| 4.05% \|`
			`\| InternLM2_7B_chat_qlora \| 37.86% \| 15.23% \| 24.34% \| 39.71% \| 22.66% \| 14.26% \| 9.21% \|`
			`\| InternLM2_7B_chat_full \| 32.45% \| 10.82% \| 20.17% \| 30.48% \| 15.67% \| 8.84% \| 5.02% \|`
Update code (#8) * feat: add agents/actions/write_markdown * [ADD] add evaluation result of base model on 5/10 epochs * Rename mother.json to mother_v1_2439.json * Add files via upload * [DOC] update README * Update requirements.txt update mpi4py installation * Update README_EN.md update English comma * Update README.md 基于母亲角色的多轮对话模型微调完毕。已上传到 Huggingface。 * 多轮对话母亲角色的微调的脚本 * Update README.md 加上了王几行XING 和思在的作者信息 * Update README_EN.md * Update README.md * Update README_EN.md * Update README_EN.md * Changes to be committed: modified: .gitignore modified: README.md modified: README_EN.md new file: assets/EmoLLM_transparent.png deleted: assets/Shusheng.jpg new file: assets/Shusheng.png new file: assets/aiwei_demo1.gif new file: assets/aiwei_demo2.gif new file: assets/aiwei_demo3.gif new file: assets/aiwei_demo4.gif * Update README.md rectify aiwei_demo.gif * Update README.md rectify aiwei_demo style * Changes to be committed: modified: README.md modified: README_EN.md * Changes to be committed: modified: README.md modified: README_EN.md * [Doc] update readme * [Doc] update readme * Update README.md * Update README_EN.md * Update README.md * Update README_EN.md * Delete datasets/mother_v1_2439.json * Rename mother_v2_3838.json to mother_v2.json * Delete datasets/mother_v2.json * Add files via upload * Update README.md * Update README_EN.md * [Doc] Update README_EN.md minor fix * InternLM2-Base-7B QLoRA微调模型链接和测评结果更新 * add download_model.py script, automatic download of model libraries * 清除图片的黑边、更新作者信息 modified: README.md new file: assets/aiwei_demo.gif deleted: assets/aiwei_demo1.gif modified: assets/aiwei_demo2.gif modified: assets/aiwei_demo3.gif modified: assets/aiwei_demo4.gif * rectify aiwei_demo transparent * transparent * modify: aiwei_demo table--->div * modified: aiwei_demo * modify: div ---> table * modified: README.md * modified: README_EN.md * update model config file links * Create internlm2_20b_chat_lora_alpaca_e3.py 20b模型的配置文件 * update model config file links update model config file links * Revert "update model config file links" --------- Co-authored-by: jujimeizuo <fengzetao.zed@foxmail.com> Co-authored-by: xzw <62385492+aJupyter@users.noreply.github.com> Co-authored-by: Zeyu Ba <72795264+ZeyuBa@users.noreply.github.com> Co-authored-by: Bryce Wang <90940753+brycewang2018@users.noreply.github.com> Co-authored-by: zealot52099 <songyan5209@163.com> Co-authored-by: HongCheng <kwchenghong@gmail.com> Co-authored-by: Yicong <yicooong@qq.com> Co-authored-by: Yicooong <54353406+Yicooong@users.noreply.github.com> Co-authored-by: aJupyter <ajupyter@163.com> Co-authored-by: MING_X <119648793+MING-ZCH@users.noreply.github.com> Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com> Co-authored-by: HatBoy <null2none@163.com> Co-authored-by: ZhouXinAo <142309012+zxazys@users.noreply.github.com> 2024-04-14 10:09:17 +08:00			`\| InternLM2_7B_base_qlora_5epoch \| 41.94% \| 20.21% \| 29.67% \| 42.98% \| 27.07% \| 19.33% \| 14.62% \|`
			`\| InternLM2_7B_base_qlora_10epoch \| 43.47% \| 22.06% \| 31.4% \| 44.81% \| 29.15% \| 21.44% \| 16.72% \|`