Downloads ยท 30 days
33
5% of all-time downloads
ShaomuTan/ReMedy-2B
ReMedy-2B is a translation model from ShaomuTan. Use it when you need text moved from one language to another. The card lists the license as apache-2.0.
Downloads ยท 30 days
33
5% of all-time downloads
All-time downloads
649
Public
Parameters
2.6B
5.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors5.2 GB ยท 99%
From the Hugging Face model README
Learning High-Quality Machine Translation Evaluation from Human Preferences with Reward Modeling
</div>ReMedy is a new state-of-the-art machine translation (MT) evaluation framework that reframes the task as reward modeling rather than direct regression. Instead of relying on noisy human scores, ReMedy learns from pairwise human preferences, leading to better alignment with human judgments.
ReMedy demonstrates that reward modeling with pairwise preferences offers a more reliable and human-aligned approach for MT evaluation.
ReMedy requires Python โฅ 3.10, and leverages VLLM for fast inference.
pip install remedy-mt-eval
git clone https://github.com/Smu-Tan/Remedy
cd Remedy
git clone https://github.com/Smu-Tan/Remedy
cd Remedy
pip install -e .
git clone https://github.com/Smu-Tan/Remedy
cd Remedy
poetry install
Python โฅ 3.10transformers โฅ 4.51.1vllm โฅ 0.8.5torch โฅ 2.6.0pyproject.toml for full dependencies)Before using, download the model from HuggingFace:
HF_HUB_ENABLE_HF_TRANSFER=1 huggingface-cli download ShaomuTan/ReMedy-9B-23 --local-dir Models/remedy-9B-23
You can replace ReMedy-9B-22 with other variants like ReMedy-9B-23.
remedy-score \
--model Models/remedy-9B-22 \
--src_file testcase/en.src \
--mt_file testcase/en-de.hyp \
--ref_file testcase/de.ref \
--src_lang en --tgt_lang de \
--cache_dir $CACHE_DIR \
--save_dir testcase \
--num_gpus 4 \
--calibrate
remedy-score \
--model Models/remedy-9B-22 \
--src_file testcase/en.src \
--mt_file testcase/en-de.hyp \
--no_ref \
--src_lang en --tgt_lang de \
--cache_dir $CACHE_DIR \
--save_dir testcase \
--num_gpus 4 \
--calibrate
src-tgt_raw_scores.txtsrc-tgt_sigmoid_scores.txtsrc-tgt_calibration_scores.txtsrc-tgt_detailed_results.tsvsrc-tgt_result.jsonInspired by SacreBLEU, ReMedy provides JSON-style results to ensure transparency and comparability.
<details> <summary>๐ Example JSON Output</summary>{
"metric_name": "remedy-9B-22",
"raw_score": 4.502863049214531,
"sigmoid_score": 0.9613502018042875,
"calibration_score": 0.9029647169507162,
"calibration_temp": 1.7999999999999998,
"signature": "metric_name:remedy-9B-22|lp:en-de|ref:yes|version:0.1.1",
"language_pair": "en-de",
"source_language": "en",
"target_language": "de",
"segments": 2037,
"version": "0.1.1",
"args": {
"src_file": "testcase/en.src",
"mt_file": "testcase/en-de.hyp",
"src_lang": "en",
"tgt_lang": "de",
"model": "Models/remedy-9B-22",
"cache_dir": "Models",
"save_dir": "testcase",
"ref_file": "testcase/de.ref",
"no_ref": false,
"calibrate": true,
"num_gpus": 4,
"num_seqs": 256,
"max_length": 4096,
"enable_truncate": false,
"version": false,
"list_languages": false
}
}
</details>
--src_file # Path to source file
--mt_file # Path to MT output file
--src_lang # Source language code
--tgt_lang # Target language code
--model # Model path or HuggingFace ID
--save_dir # Output directory
--ref_file # Reference file path
--no_ref # Reference-free mode
--cache_dir # Cache directory
--calibrate # Enable calibration
--num_gpus # Number of GPUs
--num_seqs # Number of sequences (default: 256)
--max_length # Max token length (default: 4096)
--enable_truncate # Truncate sequences
--version # Print version
--list_languages # List supported languages
</details>
| Model | Size | Base Model | Ref/QE | Download |
|---|---|---|---|---|
| ReMedy-2B | 2B | Gemma-2-2B | Both | ๐ค HuggingFace |
| ReMedy-9B-22 | 9B | Gemma-2-9B | Both | ๐ค HuggingFace |
| ReMedy-9B-23 | 9B | Gemma-2-9B | Both | ๐ค HuggingFace |
| ReMedy-9B-24 | 9B | Gemma-2-9B | Both | ๐ค HuggingFace |
More variants coming soon...
mt-metrics-evalgit clone https://github.com/google-research/mt-metrics-eval.git
cd mt-metrics-eval
pip install .
python3 -m mt_metrics_eval.mtme --download
bash wmt/wmt22.sh
bash wmt/wmt23.sh
bash wmt/wmt24.sh
</details>๐ Results will be comparable with other metrics reported in WMT shared tasks.
If you use ReMedy, please cite the following paper:
@inproceedings{tan-monz-2025-remedy,
title = "{R}e{M}edy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling",
author = "Tan, Shaomu and
Monz, Christof",
editor = "Christodoulopoulos, Christos and
Chakraborty, Tanmoy and
Rose, Carolyn and
Peng, Violet",
booktitle = "Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing",
month = nov,
year = "2025",
address = "Suzhou, China",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.emnlp-main.217/",
doi = "10.18653/v1/2025.emnlp-main.217",
pages = "4370--4387",
ISBN = "979-8-89176-332-6",
abstract = "A key challenge in MT evaluation is the inherent noise and inconsistency of human ratings. Regression-based neural metrics struggle with this noise, while prompting LLMs shows promise at system-level evaluation but performs poorly at segment level. In this work, we propose ReMedy, a novel MT metric framework that reformulates translation evaluation as a reward modeling task. Instead of regressing on imperfect human ratings directly, ReMedy learns relative translation quality using pairwise preference data, resulting in a more reliable evaluation. In extensive experiments across WMT22-24 shared tasks (39 language pairs, 111 MT systems), ReMedy achieves state-of-the-art performance at both segment- and system-level evaluation. Specifically, ReMedy-9B surpasses larger WMT winners and massive closed LLMs such as MetricX-13B, XCOMET-Ensemble, GEMBA-GPT-4, PaLM-540B, and finetuned PaLM2. Further analyses demonstrate that ReMedy delivers superior capability in detecting translation errors and evaluating low-quality translations."
}