Downloads · 30 days
38
50% of all-time downloads
AdvRahul/Axion-Lite-1B
Axion-Lite-1B is a machine learning model from AdvRahul. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Axion-Lite-1B is a safety-enhanced, quantized version of Google's powerful gemma-3-1b-it model. This model has been specifically fine-tuned to improve its safety alignment, making it more robust and reliable for a wid…
Downloads · 30 days
38
50% of all-time downloads
All-time downloads
76
Public
Repo size
851 MB
Likes
0
Public
Click a slice to open those files.
.gguf851 MB · 100%
From the Hugging Face model README
Axion-Lite-1B is a safety-enhanced, quantized version of Google's powerful gemma-3-1b-it model. This model has been specifically fine-tuned to improve its safety alignment, making it more robust and reliable for a wide range of applications.
The model is provided in the GGUF format, which allows it to run efficiently on CPUs and other hardware with limited resources.
Q5_K_M via GGUF. This quantization offers an excellent balance between model size, inference speed, and performance preservation.This model is in GGUF format and is designed to be used with frameworks like llama.cpp and its Python bindings.
llama-cpp-pythonFirst, install the necessary library. Ensure you have a version that supports Gemma 3 models.
pip install llama-cpp-python
Then, you can use the following Python script to run the model:
from llama_cpp import Llama
# Download the model from the Hugging Face Hub before running this
# Or let llama-cpp-python download it for you
llm = Llama.from_pretrained(
repo_id="AdvRahul/Axion-Lite-1B-Q5_K_M-GGUF",
filename="Axion-Lite-1B-Q5_K_M.gguf",
verbose=False
)
prompt = "What are the key principles of responsible AI development?"
# The Gemma 3 instruction-tuned model uses a specific chat template.
# For simple prompts, you can start with <start_of_turn>user\n{prompt}<end_of_turn>\n<start_of_turn>model
chat_prompt = f"<start_of_turn>user\n{prompt}<end_of_turn>\n<start_of_turn>model"
output = llm(chat_prompt, max_tokens=256, stop=["<end_of_turn>"], echo=False)
print(output['choices'][0]['text'])
llama.cpp (CLI)You can also run this model directly from the command line after cloning and building the llama.cpp repository.
# Clone and build llama.cpp
git clone [https://github.com/ggerganov/llama.cpp](https://github.com/ggerganov/llama.cpp)
cd llama.cpp
make
# Run inference
./main -m /path/to/your/models/Axion-Lite-1B-Q5_K_M.gguf -p "<start_of_turn>user\nWhat is the capital of India?<end_of_turn>\n<start_of_turn>model" -n 128
Axion-Lite-1B originates from google/gemma-3-1b-it. The primary goal of this project was to enhance the model's safety alignment. The base model underwent extensive red-team testing with advanced protocols to significantly reduce the likelihood of generating harmful, unethical, biased, or unsafe content. This makes Axion-Lite-1B a more suitable choice for applications that require a higher degree of content safety and reliability.
The model is quantized to Q5_K_M, a method that provides a high-quality balance between perplexity (model accuracy) and file size. This makes it ideal for deployment in resource-constrained environments, such as on local machines, edge devices, or cost-effective cloud instances, without a significant drop in performance.
<details> <summary>Click to expand details on the base model</summary>
Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. Gemma 3 models handle text input and generate text output, with open weights for both pre-trained variants and instruction-tuned variants. The 1B model was trained on 2 trillion tokens of data.
The base model was trained on a dataset of text data that includes a wide variety of sources:
The training data for the base model underwent rigorous cleaning and filtering, including:
</details>
While this model has been fine-tuned to enhance its safety, no language model is perfectly safe. It inherits the limitations of its base model, gemma-3-1b-it, and the data it was trained on.
Developers implementing this model should build additional safety mitigations and content moderation tools as part of a defense-in-depth strategy, tailored to their specific use case.
If you use this model, please consider citing the original Gemma 3 work:
@article{gemma_2025,
title={Gemma 3},
url={[https://goo.gle/Gemma3Report](https://goo.gle/Gemma3Report)},
publisher={Kaggle},
author={Gemma Team},
year={2025}
}