Skip to content

VanWang

DeepSeek-R1-Distill-Qwen-7B-ThinkPO

VanWang/DeepSeek-R1-Distill-Qwen-7B-ThinkPO

DeepSeek-R1-Distill-Qwen-7B-ThinkPO is a machine learning model from VanWang. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.

a simple yet effective postSFT method that enhances long CoT reasoning without requiring new long CoT responses - Here, we show the results of open-source reasoning LLMs before and after ThinkPO.

Downloads · 30 days

34

10% of all-time downloads

All-time downloads

330

Public

Parameters

7.6B

15.2 GB on disk

Likes

2

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors15.2 GB · 100%

At a glance

License
mit
Model type
qwen2
Access
Public
Created
Feb 20, 2025
Updated
Feb 21, 2025
SHA
74c97a39
Type
qwen2
License
mit
Created
Feb 20, 2025
Updated
Feb 21, 2025