Downloads · 30 days
30
10% of all-time downloads
zfj1998/AWA-RL
AWA-RL is a text generation model from zfj1998. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
AWA-RL is a search-agent checkpoint trained with Abstention-Aware Reinforcement Learning. It is designed to answer when retrieved evidence is sufficient and abstain with not enough information when the available evide…
Downloads · 30 days
30
10% of all-time downloads
All-time downloads
300
Public
Parameters
7.6B
15.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors15.2 GB · 100%
From the Hugging Face model README
AWA-RL is a search-agent checkpoint trained with Abstention-Aware Reinforcement Learning. It is designed to answer when retrieved evidence is sufficient and abstain with not enough information when the available evidence does not support a reliable answer.
The checkpoint uses the Qwen2ForCausalLM architecture and is stored in BF16. Code, evaluation scripts, and validation subsets are available in the AWA-RL GitHub repository.
AWA-RL introduces a courage factor that controls the model's willingness to answer. Adjusting this factor provides a stable way to control the refusal rate and select an operating point that balances answer coverage (recall/accuracy) against precision.

These results are from the AWA-RL paper. Across both from-scratch and cold-start training, the refusal-rate dynamics respond consistently to the courage factor. The evaluation tables show how this control enables a balance between accuracy, precision, refusal rate, and Reliability-Aware F1 (RA-F1).
The paper link will be added after the preprint is available.
The checkpoint can be loaded with Transformers:
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "zfj1998/AWA-RL"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
trust_remote_code=True,
)
For the complete search-agent loop, retrieval server setup, prompts, and evaluation commands, see the repository README. The released evaluation code expects a FlashRAG-compatible retrieval endpoint and uses vLLM for inference.
The repository includes evaluation subsets for MuSiQue, HotpotQA, and 2WikiMultiHopQA. Reported metrics are:
This is a research checkpoint intended for search-agent experiments. Its reliability depends on retrieval quality, prompting, decoding settings, and the evaluation domain. Abstention reduces unsupported answers but does not guarantee factual correctness. Users should evaluate the checkpoint for their own domain and safety requirements before deployment.
The paper citation and preprint link will be added when the paper is publicly available.