Downloads Β· 30 days
19
1% of all-time downloads
BigJuicyData/Anni
Anni is a machine learning model from BigJuicyData. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
<h1 align="center" <img src="assets/logo.png" alt="Anni Logo" width="100" / <br / Anni </h1
Downloads Β· 30 days
19
1% of all-time downloads
All-time downloads
2.2K
Public
Parameters
14.8B
29.5 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors29.5 GB Β· 100%
From the Hugging Face model README
| Property | Value |
|---|---|
| Base Model | Qwen3 14B |
| Model Type | Language Model for Code |
| Context Length | 32,000 tokens |
| Precision | BF16 / safetensors (merged) |
| Inference Framework | vLLM compatible |
Get started immediately using the provided Google Colab notebooks:
(Recommended) GGUF Inference : Open the Colab Notebook to run standard inference.
vLLM Serving: Open the Colab Notebook to run inference using the vLLM server.
pip install -r requirements.txt
tmux is installed on your system (required for training scripts).Environment Variables: Rename the example environment file and add your API tokens (WandB, HuggingFace, ModelScope).
mv config/example.env config/.env
# Edit config/.env with your keys
Training Config: Edit config/config.yaml to adjust hyperparameters.
LOCAL_STORAGE_PATH in src/train.py before starting training.To start the training process, run the shell script:
./scripts/train.sh
src/)| File | Description |
|---|---|
preprocess.py | Downloads the OpenCodeReasoning-2 dataset and preprocesses it for training. |
train.py | Downloads the base model and fine-tunes it on the preprocessed dataset. |
save.py | Loads the fine-tuned LoRA adapters and saves the model as merged 16-bit and GGUF formats. |
upload.py | Uploads the merged model to Hugging Face and ModelScope. |
scripts/)| File | Description |
|---|---|
train.sh | Runs the training script with specified parameters. |
eval.sh | Evaluates the model on the LiveCodeBench dataset. |
serve.sh | Serves the model using the vLLM server. |
terminate_train.sh | Terminates the training process. |
web/)The frontend code for Anni is available in the web directory.
π View Frontend Documentation
This repositoryβs model and its training code are released under the MIT License.
All other elements, such as frontend code, project name and logo, are trademarks of the developer and owner of this repository (Hans) and may not be used without explicit permission.
The training dataset includes openly licensed sources under CC-BY-4.0, which permits commercial use with attribution.
Attribution:
Note: The dataset itself is not included in this model release.
This model may generate incorrect or unsafe code. Evaluate and verify outputs before using in production.