Downloads · 30 days
19
17% of all-time downloads
tamewild/PCSS-Qwen3-4B
PCSS-Qwen3-4B is a text generation model from tamewild. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
GitHub Repository | Blog Post | Reproduction Notebook
Downloads · 30 days
19
17% of all-time downloads
All-time downloads
114
Public
Parameters
4B
8.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8 GB · 100%
From the Hugging Face model README
GitHub Repository | Blog Post | Reproduction Notebook
PCSS-Qwen3-4B is an experimental checkpoint demonstrating that fine-tuning small base models on pure logic deduction traces can elicit substantial out-of-domain reasoning generalization. Starting from unsloth/Qwen3-4B-Base (Unsloth's mirror of Qwen/Qwen3-4B-Base), this model was fine-tuned on 100 5x5 zebra puzzles from the tamewild/zebra_100 dataset, which contains zero mathematical data.
For fine-tuning on these long-form logic traces, the model was trained using PCSS (Per-Example Calibrated Sigmoid Scaler), an adaptive loss scaler derived from KTO. The training process leveraged MiSS (Matrix Shard Sharing) for parameter-efficient fine-tuning (PEFT), completing in approximately 6.5 minutes on a single NVIDIA H200 NVL GPU. (Note: Given the minimal 100-example training regime, independent runs exhibit higher variance than 500-sample configurations; reproduction runs generally yield MATH-500 scores in the 80%–86% range due to optimization and inference non-determinism).
[!WARNING] Experimental Checkpoint Notice:
PCSS-Qwen3-4Bis an exploratory research artifact developed solely to study out-of-domain reasoning generalization. UnlikePCSS-Qwen3.5-9B, this model has a significantly narrower capability profile and lacks general conversational robustness. It is not intended for chat, multi-turn dialogue, or general assistant tasks.
Our fine-tuned checkpoint was evaluated via vLLM (0.19.1) using greedy decoding with 16,384 max completion tokens. All reported scores are Pass@1 estimates. To account for runtime batching non-determinism in vLLM, scores for MATH-500, 4x4 Zebra, and 5x5 Zebra were averaged over 3 evaluation runs, and AIME 2025 was averaged over 6 runs. Full MATH (5,000 problems) was evaluated with a single greedy run.
Baseline scores for the untrained base model and the official post-trained model in non-thinking mode are self-reported from the official Qwen 3 Technical Report.
| Benchmark | Qwen 3 4B Base (PCSS + MiSS) | Qwen 3 4B Post-Trained (Self-Reported Non-Thinking Mode) | Qwen 3 4B Base (Self-Reported Baseline) |
|---|---|---|---|
| MATH-500 | 84.60% | 84.80% | — |
| Full MATH (5,000) | 85.26% (+31.16% delta) | — | 54.10% (4-shot) |
| AIME 2025 | 21.67% | 19.10% | — |
| 4x4 Zebra | 31.67% | — | — |
| 5x5 Zebra | 4.33% | — | — |
For full training details and mathematical derivations, please refer to the Reproduction Notebook and Blog Post.
unsloth/Qwen3-4B-Basetamewild/zebra_100 (100 5x5 zebra puzzles)The checkpoint bundles the pre-configured chat_template.jinja. To serve the model via vLLM:
vllm serve tamewild/PCSS-Qwen3-4B \
--max-model-len 32000 \
--generation-config vllm \
--host 127.0.0.1 \
--port 18000 \
--gpu-memory-utilization 0.90
Evaluations were conducted using the following zero-shot Chain-of-Thought templates:
{problem}
Please reason step by step, and put your final answer within \boxed{}
{problem}
Provide the solution grid as the final answer.
Please reason step by step, and put your final answer within \boxed{}