Downloads · 30 days
0
Yale-ROSE/Qwen3-4B-SAT-VarSelector-Sym-Aug
Qwen3-4B-SAT-VarSelector-Sym-Aug is a text classification model from Yale-ROSE. Use it when you need a label for a piece of text. The card lists the license as apache-2.0.
A Qwen3-4B model fine-tuned for SAT branching variable selection using symmetry-based data augmentation.
Downloads · 30 days
0
Access
Public
Updated Jan 21, 2026
Repo size
56.3 GB
Likes
0
Public
Click a slice to open those files.
.pt48.3 GB · 86%
From the Hugging Face model README
A Qwen3-4B model fine-tuned for SAT branching variable selection using symmetry-based data augmentation.
This model predicts which variable to branch/cube on next, given a SAT CNF formula state. It was trained with 5x augmented data using CNF symmetry transformations, achieving 21.8% top-1 accuracy (vs 19% for Qwen3-0.6B).
Qwen/Qwen3-4B (causal language model)This model was trained with 5x data augmentation using semantically-safe CNF transformations:
| Augmentation | Description | Effect |
|---|---|---|
| Variable Permutation | Bijective remapping of variable IDs | Prevents memorizing specific variable numbers |
| Clause Shuffling | Random reordering of clauses | Teaches position-independence |
| Literal Reordering | Shuffle literals within clauses | Token-level variation |
| Polarity Flipping | Flip signs of random variable subset | Teaches structural vs. polarity features |
| Parameter | Value |
|---|---|
| Original training samples | 8,110 |
| Augmented training samples | 40,550 (5x) |
| Validation samples | 902 (unaugmented) |
| Epochs | 3 |
| Hardware | 8×H100 GPUs |
| Training framework | DeepSpeed ZeRO-3 |
| Peak learning rate | 5e-6 |
| Training time | ~4 hours |
| Best checkpoint | Step 1850 (epoch 2.92) |
| Model | Parameters | Training Data | Top-1 Accuracy |
|---|---|---|---|
| Qwen3-0.6B (baseline) | 600M | 8,110 samples | ~12% |
| Qwen3-0.6B (augmented) | 600M | 40,550 samples | ~19% |
| Qwen3-4B (augmented) | 4B | 40,550 samples | ~22% |
import torch
from transformers import AutoTokenizer
from sft_qwen_var_classifier import QwenVarClassifier, cnf_valid_mask
# Load model
model = QwenVarClassifier("Qwen/Qwen3-4B", max_vars=600)
state_dict = torch.load("pytorch_model.bin", map_location="cpu")
model.load_state_dict(state_dict, strict=False)
model = model.to("cuda", dtype=torch.bfloat16)
model.eval()
# Load tokenizer
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B")
# Prepare CNF input
cnf_text = """p cnf 100 250
1 -2 3 0
-1 2 -4 0
...
"""
# Tokenize
inputs = tokenizer(cnf_text, return_tensors="pt", truncation=True, max_length=8192)
inputs = {k: v.to("cuda") for k, v in inputs.items()}
# Get valid variable mask
valid_mask = torch.tensor([cnf_valid_mask(cnf_text, max_vars=600)], dtype=torch.bool, device="cuda")
# Predict
with torch.no_grad():
outputs = model(**inputs)
logits = outputs["logits"]
logits = logits.masked_fill(~valid_mask, -1e4)
predicted_var = logits.argmax(dim=-1).item()
print(f"Predicted branching variable: {predicted_var}")
pytorch_model.bin - Model weights (~8GB, bfloat16)sft_qwen_var_classifier.py - Model class definition (required for loading)If you use this model, please cite the Transformer-CnC paper.
Apache 2.0