Downloads · 30 days
16
7% of all-time downloads
xJoePec/galena-2b-math-physics
galena-2b-math-physics is a text generation model from xJoePec. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
16
7% of all-time downloads
All-time downloads
246
Public
Repo size
10.1 GB
Likes
1
Public
Click a slice to open those files.
.gguf5.1 GB · 50%
From the Hugging Face model README
A specialized 2B parameter language model fine-tuned on advanced mathematics and physics datasets. Built on IBM's Granite 3.3-2B Instruct base model with LoRA fine-tuning on 26k instruction-response pairs covering advanced calculations and physics concepts.
The HF checkpoint and GGUF exports are hosted externally (e.g., Hugging Face) and are not stored inside this repository. Fetch them before running the examples:
python scripts/download_artifacts.py --artifact all
--source huggingface (default) pulls from xJoepec/galena-2b-math-physics.--source mirror --hf-url ... --gguf-url ... lets you point to release assets/CDN downloads instead.Artifacts install under models/math-physics/{hf,gguf} and are ignored by Git.
from transformers import AutoModelForCausalLM, AutoTokenizer
# Load model and tokenizer
model = AutoModelForCausalLM.from_pretrained(
"models/math-physics/hf",
device_map="auto",
trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained("models/math-physics/hf")
# Generate response
prompt = "Explain the relationship between energy and momentum in special relativity."
messages = [{"role": "user", "content": prompt}]
inputs = tokenizer.apply_chat_template(messages, return_tensors="pt").to(model.device)
outputs = model.generate(inputs, max_new_tokens=256, temperature=0.7)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
# Requires llama.cpp build and downloaded GGUF artifact
./llama.cpp/build/bin/llama-cli \
-m models/math-physics/gguf/granite-math-physics-f16.gguf \
-p "Calculate the escape velocity from Earth's surface." \
-n 256 \
--temp 0.7
| Format | Location (after download) | Size | Use Case |
|---|---|---|---|
| Hugging Face | models/math-physics/hf/ | ~5.0 GB | PyTorch, Transformers, vLLM, further fine-tuning |
| GGUF (F16) | models/math-physics/gguf/ | ~4.7 GB | llama.cpp, Ollama, LM Studio, on-device inference |
huggingface_hub (installed via pip install -r requirements.txt) for scripted downloads# Clone repository
git clone <repository-url>
cd galena-2B
# Install dependencies
pip install -r requirements.txt
# Download artifacts (Hugging Face by default)
python scripts/download_artifacts.py --artifact hf
# Clone llama.cpp (if not already available)
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
# Build with CUDA support (Linux/WSL)
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release
# Run inference
python scripts/download_artifacts.py --artifact gguf
./build/bin/llama-cli -m ../galena-2B/models/math-physics/gguf/granite-math-physics-f16.gguf
See the examples/ directory for detailed usage demonstrations:
The model was fine-tuned using the following configuration:
# LoRA fine-tuning
python scripts/train_lora.py \
--base_model ibm-granite/granite-3.3-2b-instruct \
--dataset_path data/math_physics.jsonl \
--output_dir outputs/granite-math-physics-lora \
--use_4bit --gradient_checkpointing \
--per_device_train_batch_size 1 \
--gradient_accumulation_steps 4 \
--num_train_epochs 1 \
--max_steps 500 \
--batching_strategy padding \
--max_seq_length 512 \
--bf16 \
--trust_remote_code
For detailed training methodology and dataset preparation, see MODEL_CARD.md.
Strengths:
Limitations:
If you use this model in your research, please cite:
@software{galena_2b_2024,
title = {Galena-2B: Granite 3.3 Math & Physics Model},
author = {Your Name},
year = {2024},
url = {https://github.com/yourusername/galena-2B},
note = {Fine-tuned from IBM Granite 3.3-2B Instruct}
}
Also cite the base Granite model:
@software{granite_3_3_2024,
title = {Granite 3.3: IBM's Open Foundation Models},
author = {IBM Research},
year = {2024},
url = {https://www.ibm.com/granite}
}
This project is licensed under the Apache License 2.0 - see the LICENSE file for details.
The base Granite 3.3 model is also released under Apache 2.0 by IBM.
For issues, questions, or contributions, please open an issue in this repository's issue tracker.