Downloads · 30 days
0
13lanko/arch-assistant-final-lora
arch-assistant-final-lora is a machine learning model from 13lanko. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
The complete source code, raw data, and training scripts for this project can be found on GitHub: https://github.com/13lanko/ArchPCAssistent
Downloads · 30 days
0
Access
Public
Updated May 10, 2026
Repo size
162 MB
Likes
0
Public
Click a slice to open those files.
.safetensors162 MB · 100%
From the Hugging Face model README
The complete source code, raw data, and training scripts for this project can be found on GitHub: https://github.com/13lanko/Arch_PC_Assistent
The goal of this project is the development of a pipeline to train smaller language models locally into specialized assistants. The focus is on a multi-stage approach using PEFT (LoRA) and RSFT (Rejection Sampling Fine-Tuning).
uv.qwen2.5-7b-instruct-unsloth-bnb-4bit.<think> tags), inspired by modern reasoning models like DeepSeek R1.BAAI/bge-small-en-v1.5.Evaluated by an API Judge (Scale 1-10) on 50 new, unseen troubleshooting questions.
| Model | Format Adherence | Tech. Correctness | Usefulness | Total Score |
|---|---|---|---|---|
| BASE | 9.7 | 4.3 | 5.0 | 6.3 |
| SFT | 8.2 | 5.3 | 5.8 | 6.4 |
| RSFT | 9.4 | 5.8 | 6.2 | 7.1 |
In summary, the primary goal of this project—developing a resource-efficient pipeline to create a local AI assistant for Arch Linux and Hyprland—was fully achieved. The approach impressively demonstrates that smaller open-source models (such as Qwen2.5-7B) can be successfully trained into highly specialized experts using consumer hardware.
The decision to forgo the extremely compute-intensive GRPO and instead opt for a multi-stage approach using Rejection Sampling Fine-Tuning (RSFT) proved to be a decisive success factor. By combining SFT for fundamental behavioral alignment and RSFT for advanced logical reasoning, coupled with direct Retrieval-Augmented Generation (RAG) using ChromaDB, the model's performance was significantly enhanced without exceeding hardware limits.
The final LLM-as-a-Judge evaluation visualizes the model's learning process:
This project demonstrates that when specializing local language models, the quality of data and a well-thought-out architecture (Knowledge Distillation, RLAIF-supported filtering, RAG integration) far outweigh sheer model size or infinite computing power.