Downloads · 30 days
469
76% of all-time downloads
tamewild/PCSS-Qwen3.5-9B
PCSS-Qwen3.5-9B is a image-text-to-text model from tamewild. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
GitHub Repository | Blog Post | Reproduction Notebook
Downloads · 30 days
469
76% of all-time downloads
All-time downloads
617
Public
Parameters
9.7B
19.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors19.3 GB · 100%
From the Hugging Face model README
GitHub Repository | Blog Post | Reproduction Notebook
PCSS-Qwen3.5-9B is an experimental checkpoint demonstrating that fine-tuning small base models on pure logic deduction traces can elicit substantial out-of-domain reasoning generalization. Starting from Qwen/Qwen3.5-9B-Base, this model was fine-tuned on just 500 5x5 zebra puzzles from the tamewild/instruct5 dataset, which contains zero mathematical data.
For fine-tuning on these long-form logic traces, the model was trained using PCSS (Per-Example Calibrated Sigmoid Scaler), an adaptive loss scaler derived from KTO. The training process leveraged MiSS (Matrix Shard Sharing) for parameter-efficient fine-tuning (PEFT) alongside a preliminary PCSS extension for Multi-Token Prediction (MTP) with k=1, completing in approximately 40 minutes on a single NVIDIA H200 NVL GPU. (Note: The accompanying reproduction notebook was run on an H100 SXM, producing comparable results but with minor variance due to training non-determinism).
Qwen/Qwen3.5-9B-Base model page, the base checkpoint was released with chat template control tokens (such as <|im_start|> and <|im_end|>) explicitly pre-trained "to allow efficient LoRA-style PEFT with the official chat template," mitigating the need to fine-tune embeddings.enable_thinking) that controls whether the model generates a chain of thought within <think> tags prior to its final response. In our experiments, we only train and evaluate in non-thinking mode (enable_thinking=False). This ensures the untrained base, our fine-tuned checkpoint, and the official post-trained model are evaluated under strictly comparable conditions without generation within these tags.Models were served via vLLM (0.19.1) using the following configuration:
0.0 for the untrained base and our fine-tuned model (1.5 was applied for the official post-trained checkpoint as recommended by Qwen).To account for sampling variance, most benchmarks were evaluated using multiple independent samples. All reported scores are Pass@1 estimates. Pass@3 estimates are only reported for benchmarks with sufficient sample sizes (n >= 6).
| Benchmark | Qwen 3.5 9B Base (PCSS + MiSS) | Qwen 3.5 9B Official Post-Trained | Qwen 3.5 9B Base (Untrained) |
|---|---|---|---|
| MMLU Redux | 90.45% | 90.15% | 89.32% |
| GPQA Diamond | 71.04% | 79.63% | 61.11% |
| MATH-500 | 96.60% | 97.40% | 93.53% |
| AIME 2025 | 60.67% (Pass@3: 75.58%) | 60.56% (Pass@3: 77.50%) | 49.44% (Pass@3: 61.83%) |
| AIME 2026 | 64.33% (Pass@3: 75.50%) | 69.44% (Pass@3: 84.67%) | 50.00% (Pass@3: 65.50%) |
| ArXivMath 05/26 | 22.92% (Pass@3: 34.88%) | 24.17% (Pass@3: 31.87%) | 5.42% (Pass@3: 12.00%) |
| 5x5 Zebra | 80.50% (Pass@3: 97.05%) | 48.00% | 17.00% |
| 6x6 Zebra | 26.00% | 5.00% | — |
For full training details and mathematical derivations, please refer to the Reproduction Notebook and Blog Post.
Qwen/Qwen3.5-9B-Basetamewild/instruct5 (500 5x5 zebra puzzles)The checkpoint bundles the pre-configured chat_template.jinja. To serve the model in non-thinking mode with native MTP speculative decoding:
# Note: Configured to 2 tokens for better throughput during our benchmarking, despite k=1 fine-tuning.
vllm serve tamewild/PCSS-Qwen3.5-9B \
--max-model-len 32000 \
--language-model-only \
--generation-config vllm \
--host 127.0.0.1 \
--port 18000 \
--gpu-memory-utilization 0.90 \
--default-chat-template-kwargs '{"enable_thinking": false}' \
--speculative-config '{"method": "mtp", "num_speculative_tokens": 2}'
All evaluations were performed without a system prompt. The following zero-shot Chain-of-Thought templates were used (note the specific divergence for zebra puzzles, where the official post-trained model was evaluated using a prompt that requested a markdown table):
{problem}
Please reason step by step, and put your final answer within \boxed{}
{problem}
Please reason step by step, and put your final letter choice within \boxed{\text{}}
For the Untrained Base and PCSS Fine-Tuned Models:
{problem}
Provide the solution grid as the final answer.
Please reason step by step, and put your final answer within \boxed{}
For the Official Post-Trained Model:
{problem}
Provide the solution table as the final answer.
Because the official post-trained model struggled to consistently output parsable LaTeX grids, its prompt was adjusted to request a markdown table. While this adjusted prompt omitted the explicit Chain-of-Thought instruction, the model natively outputs extensive step-by-step reasoning traces on these puzzles regardless. The table below contextualizes the generation lengths (in tokens) alongside benchmark accuracy across the 5x5 and 6x6 zebra puzzles.
| Model | Task | Accuracy | Median Tokens | Average Tokens |
|---|---|---|---|---|
| Qwen 3.5 9B Official Post-Trained | 5x5 Zebra | 48.00% | 11,580.0 | 11,367.3 |
| Qwen 3.5 9B Base (PCSS Fine-Tuned) | 5x5 Zebra | 80.50% | 9,209.5 | 10,126.1 |
| Qwen 3.5 9B Official Post-Trained | 6x6 Zebra | 5.00% | 13,773.0 | 13,772.0 |
| Qwen 3.5 9B Base (PCSS Fine-Tuned) | 6x6 Zebra | 26.00% | 15,543.5 | 14,654.2 |