Downloads · 30 days
158
11% of all-time downloads
VertexAGI/amethyst-1-mini
amethyst-1-mini is a text generation model from VertexAGI. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as gemma.
Amethyst 1 Mini is a general-purpose chat and instruction-following model, fine-tuned from Gemma 3 4B IT using LoRA on a small, high-quality distilled instruction dataset. It is the first model in the Amethyst family…
Downloads · 30 days
158
11% of all-time downloads
All-time downloads
1.5K
Public
Parameters
4.6B
12 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors9.1 GB · 76%
From the Hugging Face model README
Amethyst 1 Mini is a general-purpose chat and instruction-following model, fine-tuned from Gemma 3 4B IT using LoRA on a small, high-quality distilled instruction dataset. It is the first model in the Amethyst family — an early, small-scale release focused on validating an end-to-end distillation → fine-tune → evaluate pipeline on consumer hardware.
| Developed by | Independent research project |
| Base model | google/gemma-3-4b-it |
| Fine-tuning base checkpoint | mlx-community/gemma-3-4b-it-qat-4bit |
| Architecture | Gemma 3, 4B parameters (dense, decoder-only transformer) |
| Fine-tuning method | LoRA (rank 8, scale 20.0), fused into the base weights and dequantized to fp16 for this release |
| Fine-tuning framework | MLX / mlx-lm, on Apple Silicon |
| Trained modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj across 16 layers |
| Language | English |
| License | Gemma Terms of Use |
Amethyst 1 Mini was fine-tuned on 1,122 instruction/response pairs (1,082 train / 40 validation), synthetically generated via knowledge distillation from nvidia/nemotron-3-super-120b-a12b (Nemotron-3-Super, a 120B-parameter MoE model, ~12B active) through the OpenRouter API.
The dataset spans a deliberately broad set of general-chat categories to encourage well-rounded conversational ability rather than narrow task performance:
Each example was generated from a unique prompt, distilled from the teacher model, then filtered/cleaned before training.
This release merges the selected LoRA adapter into the base model and dequantizes the result to fp16, so it can be loaded directly with transformers without any MLX or quantization dependencies.
Amethyst 1 Mini is intended as a lightweight, general-purpose conversational assistant — for experimentation, research into small-scale distillation pipelines, and hobbyist deployment. It is not intended for high-stakes, safety-critical, or production use.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "VertexAIco/amethyst-1-mini"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
messages = [{"role": "user", "content": "Explain how vaccines train the immune system, in simple terms."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))
This repo includes both:
| Format | File | Notes |
|---|---|---|
| Full precision (fp16) | model-*.safetensors | For transformers — see usage above |
| GGUF (Q4_K_M) | amethyst_1_mini_Q4_K_M.gguf | For llama.cpp and compatible runtimes (LM Studio, Ollama, etc.) |
llama-cli -hf VertexAIco/amethyst-1-mini -m amethyst_1_mini_Q4_K_M.gguf -p "Explain how vaccines train the immune system, in simple terms."
If you reference this model, please cite it as:
@misc{amethyst1mini,
title = {Amethyst 1 Mini},
author = {Independent research project},
year = {2026},
note = {LoRA fine-tune of Gemma 3 4B IT, distilled from Nemotron-3-Super-120B-A12B}
}
This model is built on Gemma and subject to the Gemma Terms of Use.