Downloads · 30 days
21
13% of all-time downloads
shivs28/jee_nujan_mix_v2_base
jee_nujan_mix_v2_base is a text generation model from shivs28. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
This is the base merged model for JEE mathematics problem solving, created by combining three specialized models using linear interpolation. This model serves as the foundation for further fine-tuning on mathematical…
Downloads · 30 days
21
13% of all-time downloads
All-time downloads
158
Public
Parameters
1.8B
7.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors7.1 GB · 100%
From the Hugging Face model README
This is the base merged model for JEE mathematics problem solving, created by combining three specialized models using linear interpolation. This model serves as the foundation for further fine-tuning on mathematical datasets.
Merged Models:
Merge Method: Linear interpolation with weight normalization Output Format: Float16 for efficiency Tokenizer: Based on DeepSeek-R1-Distill-Qwen-1.5B
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
# Load the merged base model
tokenizer = AutoTokenizer.from_pretrained("shivs28/jee_nujan_mix_v2_base")
model = AutoModelForCausalLM.from_pretrained(
"shivs28/jee_nujan_mix_v2_base",
torch_dtype=torch.float16,
device_map="auto"
)
# Use for mathematical reasoning
prompt = "Solve: What is the derivative of x^2 + 3x + 1?"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_length=200, temperature=0.7)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
This base model is intended to be:
This base model will be fine-tuned on comprehensive mathematical datasets including:
Created by the JEE NUJAN MIX team for educational purposes.
Please cite the original base models:
This model is part of the NUJAN educational initiative.