Downloads · 30 days
19
12% of all-time downloads
sigridjineth/HyperCLOVAX-SEED-Think-DeepConf-14B
HyperCLOVAX-SEED-Think-DeepConf-14B is a text generation model from sigridjineth. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
This model enhances a user-forked version of HyperCLOVAX-SEED-Think-14B by integrating ideas from Meta AI × UCSD's DeepConf. It performs confidence-based quality estimation and adaptive sampling to improve both accura…
Downloads · 30 days
19
12% of all-time downloads
All-time downloads
165
Public
Parameters
14.7B
59 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors59 GB · 100%
From the Hugging Face model README
This model enhances a user-forked version of HyperCLOVAX-SEED-Think-14B by integrating ideas from Meta AI × UCSD's DeepConf. It performs confidence-based quality estimation and adaptive sampling to improve both accuracy and efficiency.
The core of this method is the Lowest Group Confidence (LGC), a metric that uses a sliding window to identify the "most uncertain segment" of a generation path. This allows for intelligent offline filtering (Top-p% Filtering, Confidence-Weighted Voting) and online optimization (Early Abort), ultimately achieving higher accuracy at a lower computational cost.
While Self-Consistency—generating multiple paths and taking a majority vote—can improve performance on reasoning tasks, its practical application is limited by prohibitive computational costs and the noise introduced by low-quality generation paths.
The DeepConf framework addresses this by reading the model's internal token generation probability distribution (confidence) to estimate the quality of a path in real time. Simple average confidence can be misleading due to the "pitfall of averages." We instead use the sliding-window LGC metric to quantify the path's weakest link.
LGC is calculated by moving a window of size $W$ (e.g., 2048 tokens) across the entire generation path, calculating the average confidence within each window, and taking the minimum value as the quality score for the entire trajectory.
The formula is: $$\text{LGC}(\text{trajectory}) = \min_{t} \frac{1}{W}\sum_{i=t}^{t+W-1} \text{conf}(y_i)$$
Here, $\text{conf}(y_i)$ is the generation probability of token $y_i$. Our implementation defaults to using the softmax probability of the top-1 token.
We leverage the model's ChatML structure, which separates the thinking (exploration) and answer (formal response) stages, by applying a dual-threshold system: $\tau_{\text{think}} < \tau_{\text{answer}}$.
| Name | Description | Default Value (Example) |
|---|---|---|
W | Sliding window length (tokens) | 2048 |
p | Percentage for Top-p% Filtering | 10 |
M | Number of warm-up paths for calibration | 16 |
| $\tau_{\text{think}}$ | Early abort threshold for the thinking stage | Dynamic (based on warm-up) |
| $\tau_{\text{answer}}$ | Early abort threshold for the answer stage | Dynamic (based on warm-up, stricter) |
N_max | Max number of paths to sample (online) | Optional limit (e.g., 64) |
deepconf vs. originalScoring: Correct = 1, Incorrect / No Format = 0. "No Format" is treated as not attempted.
| Metric | original | deepconf | Notes |
|---|---|---|---|
| Total Correct | 8 | 10 | +2 questions correct |
| Accuracy (out of 30) | 26.7% | 33.3% | +6.7%p improvement |
| Attempts (Format OK) | 8 | 11 | deepconf attempted 3 more questions |
| Format Failures | 22 | 19 | deepconf shows better format stability |
| Head-to-Head | — | — | 2 Wins / 0 Losses / 28 Ties for deepconf |
Breakdown by Part:
original solved 2/15, while deepconf solved 4/15. The performance gain was concentrated in the more difficult second half.Note: The high number of "Format Failures" in this slice indicates that the ability to adhere to strict output formatting was a significant factor in the final score.
| Metric | Improvement with deepconf |
|---|---|
| Majority-Vote Accuracy | +20.0%p |
| Avg. Generated Tokens | –29.6% |
| Avg. Generation Time | –41.6% |
Caution: These results are based on a very small sample size (N≈10). However, they signal a meaningful improvement across accuracy, speed, and cost.
This model is ideal for mathematical and logical reasoning tasks where it offers significant sample savings and improved reliability compared to standard self-consistency.
Recommended Pipeline:
p=10 as a starting point) to the remaining high-quality paths.