Downloads · 30 days
0
Gavin-Wang/granite-abstract
granite-abstract is a machine learning model from Gavin-Wang. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
A language model with learnable continuous embeddings for abstract reasoning, trained SFT and reinforcement learning.
Downloads · 30 days
0
Access
Public
Updated Jan 7, 2026
Repo size
106 MB
Likes
2
Public
Click a slice to open those files.
.pt105 MB · 99%
From the Hugging Face model README
A language model with learnable continuous embeddings for abstract reasoning, trained SFT and reinforcement learning.
Granite-Abstract extends Granite with a continuous head that generates soft embeddings instead of discrete tokens. This enables the model to perform internal reasoning more efficiently while maintaining compatibility with standard generation.
Key characteristics:
The model operates with two parallel heads:
Input → Embedding Layer → Granite Backbone → Hidden State
↓
┌─────────────┴─────────────┐
↓ ↓
Continuous Head LM Head
(soft embeddings) (discrete tokens)
Phase 1: Abstract Mode - Model generates continuous embeddings for internal reasoning
cont_logits = continuous_head(hidden_state)
top_logits, top_indices = topk(cont_logits, 256)
next_embedding = softmax(top_logits) @ embed_layer[top_indices]
Phase 2: Transition - At </think> token, switches to output mode
Phase 3: Natural Mode - Generates discrete tokens for human-readable output
logits = lm_head(hidden_state)
next_token = argmax(softmax(logits))
<think>...</think>Answer format
Results on three standard benchmarks (1024 samples each):
| Model | MMLU | GSM8K | DROP | Overall |
|---|---|---|---|---|
| Gemma 3 (270M) | 19.24% | 0.39% | 0.98% | 6.87% |
| Granite 4 Nano (350M) | 7.13% | 4.00% | 9.96% | 7.03% |
| Granite 4 Nano SFT (350M) | 6.17% | 6.54% | 10.06% | 7.59% |
| Abstract Granite RL (500M) | 5.76% | 7.42% | 10.45% | 7.876% |
| DeepSeek R1 Qwen (1.5B) | 0.20% | 0.49% | 0.10% | 0.26% |

Despite having fewer parameters than larger baselines, Abstract Granite achieves competitive performance through specialized training for reasoning tasks.
pip install torch transformers
from abstract_model import AbstractModel
model = AbstractModel.load_from_directory(
output_dir='./models/granite-abstract',
sft_model_path='./models/granite-sft',
device='cuda'
)
messages = [{"role": "user", "content": "What is 5 + 3?"}]
formatted = model.tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
input_ids = model.tokenizer(formatted, return_tensors='pt')['input_ids'].to('cuda')
result = model.forward(input_ids, max_length=256, temperature=0.7)
response = model.tokenizer.decode(result['generated_tokens'])
print(response)
print("Modes:", result['mode_sequence']) # A=abstract, N=natural, T=transition