Downloads · 30 days
39
4% of all-time downloads
Ayansk11/FinSenti-DeepSeek-R1-1.5B
FinSenti-DeepSeek-R1-1.5B is a text generation model from Ayansk11. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
FinSenti-DeepSeek-R1-1.5B is a 1.5B-parameter model fine-tuned to read short financial text (headlines, earnings snippets, market commentary) and explain its read of them before settling on positive, negative, or neut…
Downloads · 30 days
39
4% of all-time downloads
All-time downloads
911
Public
Parameters
1.8B
3.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.6 GB · 100%
From the Hugging Face model README
FinSenti-DeepSeek-R1-1.5B is a 1.5B-parameter model fine-tuned to read short financial text (headlines, earnings snippets, market commentary) and explain its read of them before settling on positive, negative, or neutral. It's built on DeepSeek's R1 distillation, so it already had a decent reasoning prior coming in. The SFT + GRPO passes narrow that reasoning to financial sentiment specifically.
The model is part of the FinSenti collection, a scaling study of small models trained on the same data with the same recipe.
<reasoning>...</reasoning><answer>...</answer> output
format that's easy to parse downstreamIt was trained on news-style headlines and earnings snippets in English, so that's where it shines. Outside that domain you'll see the format hold up but the labels get noisier.
Two-stage recipe, same across the whole FinSenti family:
Trainer stack: Unsloth + TRL, using Unsloth's pre-quantized mirror
unsloth/DeepSeek-R1-Distill-Qwen-1.5B as the
loading shortcut for the upstream
deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B
weights. LoRA adapters (r=32, alpha=64) were
trained on the attention and MLP projection layers, then merged into the
base weights before export, so this repo is a self-contained model and
doesn't need PEFT to load.
Standard transformers usage:
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "Ayansk11/FinSenti-DeepSeek-R1-1.5B"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id, torch_dtype=torch.bfloat16, device_map="auto"
)
system = (
"You are a financial sentiment analyst. For each headline you receive, "
"write a short reasoning chain inside <reasoning>...</reasoning> tags, "
"then give a single label inside <answer>...</answer> tags. The label "
"must be exactly one of: positive, negative, neutral."
)
user = "Apple beats Q4 estimates as iPhone sales jump 12% year over year."
messages = [
{"role": "system", "content": system},
{"role": "user", "content": user},
]
prompt = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(tok.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
Expected output (your reasoning text will vary; the label should match):
<reasoning>
Beating estimates is a positive earnings surprise. A 12% YoY iPhone sales jump in the company's biggest product line points to demand strength. Both signals push the read positive.
</reasoning>
<answer>positive</answer>
The model expects the system prompt above, verbatim is best. The user turn
is the headline or short snippet you want classified. Output is two XML-ish
blocks in this order: <reasoning>...</reasoning> then
<answer>...</answer>. The <answer> content is one of positive,
negative, or neutral (lowercase, no punctuation).
If you want labels only and don't care about the reasoning, you can stop
generation as soon as you see </answer> to save tokens.
The training reward (max 4.0) hit 3.13 on the held-out validation slice. That breaks down across the four reward functions roughly as:
<reasoning> and <answer> tagsNumbers on standard finance benchmarks (FPB, FiQA, Twitter Financial News) are forthcoming and will be added once the eval pipeline lands.
bf16 weights are about 3.0 GB. You want ~4 GB of VRAM for batch=1 inference. CPU works but is slower; the Q4_K_M GGUF is the right pick if you don't have a GPU.
A few things this model isn't built for:
| Upstream base model | deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B |
| Loading mirror | unsloth/DeepSeek-R1-Distill-Qwen-1.5B (Unsloth's pre-quantized copy) |
| Dataset | Ayansk11/FinSenti-Dataset (~15.2K train per stage, 50.8K total across splits) |
| SFT length | ~0.6 hours on A100 80GB |
| GRPO budget | 3000 steps with early stopping (best near step ~360) |
| Best GRPO reward | ~3.13 / 4.0 |
| Adapter | LoRA (r=32, alpha=64) on q/k/v/o/gate/up/down projections |
| Sequence length | 2048 |
| Optimizer | AdamW (8-bit), cosine LR schedule |
| Hardware | NVIDIA A100 80GB (Indiana University BigRed200 cluster) |
| Frameworks | Unsloth + TRL |
Other sizes and bases trained with the same recipe:
There's a GGUF build of this same model at Ayansk11/FinSenti-DeepSeek-R1-1.5B-GGUF for Ollama and llama.cpp, and the dataset itself is at Ayansk11/FinSenti-Dataset.
If you're picking a size, a rough guide:
If you use this model in research, please cite:
@misc{shaikh2026finsenti,
title = {FinSenti: Small Language Models for Financial Sentiment with Chain-of-Thought Reasoning},
author = {Shaikh, Ayan},
year = {2026},
url = {https://huggingface.co/collections/Ayansk11/finsenti},
note = {Indiana University}
}
Apache 2.0, same as the base model.
Trained on the Indiana University BigRed200 cluster. Thanks to the Unsloth and TRL teams for the trainer stack, and to the Qwen / DeepSeek teams for the base models.