Downloads ยท 30 days
897
100% of all-time downloads
muradil211/AetherSearch_DPO
AetherSearch_DPO is a text generation model from muradil211. Use it when you need the model to write or continue text. It is set up for transformers.
Downloads ยท 30 days
897
100% of all-time downloads
All-time downloads
897
Public
Parameters
3.1B
6.2 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors6.2 GB ยท 100%
From the Hugging Face model README
Built on AetherSearch SFT and aligned with Direct Preference Optimization (DPO) over 2,126 search-trajectory preference pairs.
<p> <img src="https://img.shields.io/badge/Method-DPO-0F766E?style=flat-square" alt="Training method: DPO"> <img src="https://img.shields.io/badge/Pairs-2%2C126-F59E0B?style=flat-square" alt="Training pairs: 2,126"> <img src="https://img.shields.io/badge/Context-32K-2563EB?style=flat-square" alt="Context window: 32K"> </p>๐ Project ยท ๐ง SFT checkpoint
</div>๐ Bring your own retriever. AetherSearch DPO is a search-agent policy, not a self-contained question-answering service. The host runtime must execute each
<search>...</search>request and return evidence inside<information>...</information>.
Question
โ
โผ
<think>reason about what is missing</think>
โ
โผ
<search>focused retrieval query</search> โโโโโโบ Search / RAG backend
โฒ โ
โโโโโ <information>retrieved evidence</information> โโโโโโ
โ
โโโ repeat the search loop when more evidence is needed
โผ
<answer>evidence-grounded final answer</answer>
The model produces reasoning, search, and answer spans. The surrounding
runtime parses each completed <search> span, runs retrieval, appends the
result as <information>, and resumes generation until the model emits an
<answer> span.
| Field | Value |
|---|---|
| ๐งฌ Base checkpoint | muradil211/AetherSearch_SFT |
| ๐งฌ Base revision | 437aca474d3966e57e82af565db95d0ad64aa24d |
| ๐๏ธ Architecture | Qwen2 causal language model |
| ๐ข Parameters | 3,085,938,688 |
| ๐๏ธ Weight dtype | BF16 |
| ๐ Context window | 32,768 positions; training sequences capped at 4,096 |
| ๐ฏ Alignment method | Direct Preference Optimization |
Supervised fine-tuning teaches the model how to follow the search protocol; DPO then teaches it which of two valid-looking continuations is preferable. For a shared prompt \(x\), preferred continuation \(y_w\), rejected continuation \(y_l\), policy \(\pi_\theta\), and frozen SFT reference \(\pi_{\mathrm{ref}}\), training minimizes:
\log \frac{\pi_\theta(y_l \mid x)}{\pi_{\mathrm{ref}}(y_l \mid x)} \right] \right). $$
In plain terms, the policy learns to widen the preference margin between the chosen and rejected search trajectories while the frozen SFT model anchors the update. This directly optimizes pairwise preferences without training a separate reward model or running an online RL loop.
The loss is adapted to the agent-environment boundary:
<information>...</information> spans inside either
continuation are masked while remaining visible as context.<|im_end|> token;
search-terminal continuations stop before it so the runtime can insert the
next retrieval result.Both the initial policy and frozen reference use the pinned AetherSearch SFT checkpoint. The released run uses \(\beta=0.1\).
| Setting | Value | Setting | Value |
|---|---|---|---|
| Epochs | 1 | Learning rate | 5e-7 |
| DPO beta | 0.1 | Scheduler | Cosine |
| Global batch size | 12 pairs | Per-device batch | 1 pair |
| Precision | BF16 | Max sequence length | 4,096 |
| Warmup ratio | 0.03 | Weight decay | 0.0 |
| Distributed optimizer | DeepSpeed ZeRO-3 | Seed | 42 |
The policy and frozen reference are both sharded with ZeRO-3; gradients and optimizer updates are applied only to the policy.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "muradil211/AetherSearch_DPO"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
model.config.use_cache = True
Loading the checkpoint is only the first step. For end-to-end use, stop
generation after each complete <search> request, execute it with your
retriever, append the result as <information>, and resume generation. Stop
when the model emits a complete <answer> span, and preserve the XML protocol
exactly throughout the loop.
No additional blanket license is asserted here. Review the Qwen2.5-3B-Instruct license, the AetherSearch SFT terms, and the AetherSearch DPO data attribution before redistribution or downstream use.
Built for agentic search and retrieval-augmented reasoning. ๐โจ
</div>