Downloads · 30 days
6
38% of all-time downloads
INNOCUITY/DatapressoRM_Lora_v1
DatapressoRM_Lora_v1 is a machine learning model from INNOCUITY. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
The LLM agent generates trajectories using the following XML structure:
Downloads · 30 days
6
38% of all-time downloads
All-time downloads
16
Public
Parameters
8.2B
32.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors32.8 GB · 100%
From the Hugging Face model README
The LLM agent generates trajectories using the following XML structure:
<think>reasoning content</think>
<tool_call>{"name": "tool_name", "parameters": {...}}</tool_call>
<result>execution result</result>
...
<think>final reasoning</think>
<answer>final answer</answer>
<think> + <tool_call> + <result> sequence forms one clip<think> + <answer> sequence for the final outputdeepsearch: Information search toolsmicrosandbox: Code execution tools (sandbox_start, sandbox_stop, sandbox_run_code, etc.)tavily: Web crawling tools (tavily-search, tavily-extract, etc.)perform_web_task: Web interaction tools (search_google, go_to_url, etc.)Reasonableness Score (0.0-1.0) for regular clips:
Correctness Score (0.0-1.0) for final clips:
Each category has specific metrics (all 0.0-1.0):
DeepSearch: Information Relevance, Tool Use Quality, Source Quality, Information Synthesis MicroSandbox: Code Correctness, Tool Use Quality, Computational Efficiency, Result Interpretation Perform Web Task: Tool Use Quality, Content Extraction, Interaction Quality, Goal Achievement Tavily: Tool Use Quality, Information Relevance, Content Extraction, Goal Achievement Final Assessment: Task Completion, Tool Use Quality, Reasoning Coherence, Problem Resolution, Answer Correctness
This scheme focuses solely on the final output quality without considering the process.就是tool-star 那一套
将PRM 和 Output-Reward 结合
This is the core scheme that combines process evaluation with output evaluation, providing a balanced signal that values both the reasoning journey and the final destination.
The combined reward scheme is mathematically formulated as:
$$R_{final} = \begin{cases} -1.0 & \text{if } \text{format_invalid}(T) \ \alpha \cdot R_{output} + \beta \cdot R_{process} & \text{otherwise} \end{cases}$$
Where:
$$R_{output} = \begin{cases} 1.0 & \text{if } \text{format_valid}(T) \land \text{answer_correct}(T, G) \ 0.0 & \text{if } \text{format_valid}(T) \land \neg \text{answer_correct}(T, G) \ -1.0 & \text{if } \neg \text{format_valid}(T) \end{cases}$$
Where:
The process reward integrates both clip-level and category-level evaluations:
$$R_{process} = \gamma \cdot R_{clips} + (1-\gamma) \cdot R_{categories}$$
Where $\gamma = 0.6$ balances clip and category contributions.
clip score 主要注重推理的逻辑性是否合理,category score 主要注重工具调用的质量如何
Clip-Level Process Reward: $$R_{clips} = \frac{\sum_{i=1}^{n} w_i \cdot s_i}{\sum_{i=1}^{n} w_i}$$
Where:
Category-Level Process Reward: $$R_{categories} = \frac{\sum_{k=1}^{m} \lambda_k \cdot \bar{s}k}{\sum{k=1}^{m} \lambda_k}$$
Where:
The progressive weighting scheme $w_i = \frac{i}{n}$ ensures that:
This reflects the natural importance hierarchy in multi-step reasoning tasks.
This ensures that format violations receive immediate negative rewards regardless of other factors.
def apply_format_penalty_override(base_reward: float, output_format_valid: bool) -> float:
"""
Override any positive reward if format is invalid
"""
if not output_format_valid:
return -1.0
return base_reward