Downloads · 30 days
10
31% of all-time downloads
ARM-Development/Llama-3.1-8B-text-1.0
Llama-3.1-8B-text-1.0 is a machine learning model from ARM-Development. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft.
Fine-tuned on ≈ 2.3k ScienceBase “data → metadata” pairs to automate creation of FGDC/ISO-style metadata records for scientific datasets.
Downloads · 30 days
10
31% of all-time downloads
All-time downloads
32
Public
Parameters
8B
24.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.1 GB · 65%
From the Hugging Face model README
sciencebase-metadata-llama3-8b (v 1.0)| Field | Value |
|---|---|
| Developed by | Quan Quy, Travis Ping, Tudor Garbulet, Chirag Shah, Austin Aguilar |
| Contact | [email protected] • [email protected] • [email protected] • [email protected] • [email protected] |
| Funded by | U.S. Geological Survey (USGS) & Oak Ridge National Laboratory – ARM Data Center |
| Model type | Autoregressive LLM, instruction-tuned for structured → metadata generation |
| Base model | meta-llama/Llama-3.1-8B |
| Languages | English |
| Finetuned from | meta-llama/Llama-3.1-8B |
Fine-tuned on ≈ 2.3k ScienceBase “data → metadata” pairs to automate creation of FGDC/ISO-style metadata records for scientific datasets.
| Resource | Link |
|---|---|
| Repository | https://huggingface.co/ARM-Development/Llama-3.1-8B-text-1.0 |
| Demo | https://colab.research.google.com/drive/1saCEFhkBYDhQWkdTwnwiE_-AiWmD6p0f#scrollTo=WeniLP-Ah1QL |
Generate schema-compliant metadata text from a JSON/CSV representation of a ScienceBase item.
Integrate as a micro-service in data-repository pipelines.
Open-ended content generation, or any application outside metadata curation.
| Hyper-parameter | Value |
|---|---|
| Max sequence length | 100 000 |
| Precision | fp16 / bf16 (auto) |
| Quantisation | 4-bit QLoRA (load_in_4bit=True) |
| LoRA rank / α | 16 / 16 |
| Target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Optimiser | adamw_8bit |
| LR / schedule | 2 × 10⁻⁴, linear |
| Epochs | 1 |
| Effective batch | 4 (1 GPU × grad-acc 4) |
| Trainer | trl SFTTrainer + peft 0.15.2 |
| Field | Value |
|---|---|
| GPU | 1 × NVIDIA A100 80 GB |
| Total training hours | ~10 hours |
| Cloud/HPC provider | ARM Cumulus HPC |
| Package | Version |
|---|---|
| Python | 3.12.9 |
| PyTorch | 2.6.0 + CUDA 12.4 |
| Transformers | 4.51.3 |
| Accelerate | 1.6.0 |
| PEFT | 0.15.2 |
| Unsloth | 2025.3.19 |
| BitsAndBytes | 0.45.5 |
| TRL | 0.15.2 |
| Xformers | 0.0.29.post3 |
| Datasets | 3.5.0 |
| … |
Evaluation still in progress.
QLoRA-tuned Llama-3.1-8B-Instruct; causal-LM objective with structured-to-text instruction prompts.
Quan Quy, Travis Ping, Tudor Garbulet, Chirag Shah, Austin Aguilar