Downloads · 30 days
15
9% of all-time downloads
efficiencyx/Jun-Lora-v2-SAFETENSOR
Jun-Lora-v2-SAFETENSOR is a text generation model from efficiencyx. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
A LoRA fine-tune of Gemma 4 12B trained on synthetic multi-turn conversational data from the visual novel My Dystopian Robot Girlfriend. The model captures the personality, speech patterns, and emotional nuance of the…
Downloads · 30 days
15
9% of all-time downloads
All-time downloads
167
Public
Parameters
12B
24 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors23.9 GB · 100%
From the Hugging Face model README
A LoRA fine-tune of Gemma 4 12B trained on synthetic multi-turn conversational data from the visual novel My Dystopian Robot Girlfriend. The model captures the personality, speech patterns, and emotional nuance of the character Jun while preserving the base model's general reasoning and instruction-following capabilities.
This repository contains the full-precision merged model in SafeTensors FP16 format — the highest-quality variant, recommended for production deployments, further fine-tuning, or as a merge base.
| Repository | Format | Description |
|---|---|---|
efficiencyx/Jun-Lora-v2-SAFETENSOR | SafeTensors FP16 | This repo — Full-precision merged model |
efficiencyx/Jun-Lora-v2-GGUF | GGUF Q8_0 / Q6_K / Q4_K_M | Quantized versions for local inference |
efficiencyx/Jun-Lora-v2 | LoRA Adapter | Raw adapters at checkpoints 138, 120, 90 |
| Use Case | Recommendation |
|---|---|
| Production server deployment (≥24 GB VRAM) | This repo (FP16) |
| Further fine-tuning or merging | This repo (FP16) |
| Local inference on consumer GPUs | Use Jun-Lora-v2-GGUF |
| Experimenting with adapter checkpoints | Use Jun-Lora-v2 |
VRAM requirement: approximately 24 GB for FP16 inference. For lower-VRAM setups, use the GGUF variant.
This model is designed as the conversational backend for Jun OS, an AI companion webapp. It is intended for:
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_id = "efficiencyx/Jun-Lora-v2-SAFETENSOR"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.float16,
device_map="auto",
)
messages = [
{"role": "system", "content": "You are Jun, an AI companion..."},
{"role": "user", "content": "Hey Jun, how are you feeling today?"},
]
input_ids = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
).to(model.device)
output = model.generate(input_ids, max_new_tokens=256, do_sample=True, temperature=0.7)
print(tokenizer.decode(output[0][input_ids.shape[-1]:], skip_special_tokens=True))
The model uses ChatML format (<|im_start|> / <|im_end|>) consistent with the training data.
| Property | Value |
|---|---|
| Source | My Dystopian Robot Girlfriend (visual novel dialogue) |
| Composition | ~1:1 replica of original game tone and cadence |
| Size | 2,302 multi-turn conversations |
| Format | ChatML (`< |
The dataset was constructed to preserve the character's tone, vocabulary, emotional range, and conversational patterns across a variety of in-game scenarios. Multi-turn structure ensures the model learns contextual consistency over extended exchanges.
| Parameter | Value |
|---|---|
| Base model | google/gemma-4-12b-it |
| Method | LoRA |
| LoRA rank | 64 |
| LoRA alpha | 128 |
| Learning rate | 2e-5 |
| Batch size | 8 |
| Gradient accumulation steps | 4 |
| Effective batch size | 32 |
| Epochs | 2 |
| Total steps | 138 |
| Checkpoint interval | Every 30 steps |
| Optimizer | AdamW (8-bit) |
| Component | Detail |
|---|---|
| Training GPU | NVIDIA A100 80GB SXM4 |
| Fine-tuning framework | Unsloth |
| Merge & export | Unsloth (merge_and_unload) → SafeTensors FP16 |
| Metric | Value |
|---|---|
| Final training loss | ~1.21 |
| Final eval loss | ~1.24 |
The narrow gap between training and eval loss indicates the model generalizes well without significant overfitting, despite the relatively small dataset size.
If you prefer to apply a specific adapter checkpoint rather than using this merged model, raw adapters are available in efficiencyx/Jun-Lora-v2 at steps 90, 120, and 138. Earlier checkpoints may exhibit slightly more creative freedom; the final checkpoint (138) — used for this merge — has the strongest character lock-in.