Downloads · 30 days
670
7% of all-time downloads
Nikity/lille-130m-instruct
lille-130m-instruct is a text generation model from Nikity. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Downloads · 30 days
670
7% of all-time downloads
All-time downloads
10.2K
Public
Parameters
127M
3.2 GB on disk
Likes
117
Trending 1
Click a slice to open those files.
.pt1.5 GB · 47%
From the Hugging Face model README

You are currently viewing the
lille-130m-instructmodel card.View the base model here: Nikity/lille-130m-base
Lille is a 130-million-parameter language model built from the ground up as a core component of a completely open-source deep learning stack. The name Lille reflects both its compact size and strong capabilities - capturing the idea that less can be more. It draws on the Norwegian word lille (‘small’ or ‘little’) as well as the French city Lille, giving it both meaning and place. It was trained using a custom tokenizer, a curated dataset, and a memory-efficient optimizer, all of which are publicly available.
The model comes in two versions:
Lille-130M-Base: The foundational model pretrained on 4.27 billion of tokens from the FineWeb-Edu dataset. A post-processing step to only include the highest quality of content was added. It has strong general knowledge and text completion abilities.Lille-130M-Instruct: The instruction-tuned version, fine-tuned on the Kyoto-Corpus. It excels at following user commands, engaging in chat, and performing a variety of instruction-based tasks.The model architecture is a modern Transformer decoder featuring Grouped-Query Attention (GQA), RoPE, and RMSNorm, making it efficient and performant for its size.
Note on parameter count: While the model name is 130M for simplicity, the actual parameter count is 127.17 million.
All evaluations were conducted using simple-eval, our open-source evaluation framework. Benchmarks are run in a zero-shot setting unless specified otherwise.
Lille-130M-Instruct
Evaluations for other LLMs are sourced from the <a href="https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard">Open LLM Leaderboard</a> or their respective model cards when benchmark data is unavailable. For Lille 130M Instruct, evaluations are performed using <a href="https://github.com/Nikityyy/simple-eval">simple-eval</a>. ARC-C and ARC-E for Smollm2 are also evaluated using <a href="https://github.com/Nikityyy/simple-eval">simple-eval</a>.
There are several ways to use the Lille models, from easy-to-use graphical interfaces to advanced programmatic control.
LM Studio provides a simple graphical interface to run LLMs on your local machine. It's the easiest way to start chatting with Lille.
Lille or Nikity. You will find the models I have uploaded.lille-130m-instruct-f16.gguf.The easiest way to use Lille programmatically is with the simpleai-sdk, which handles all the boilerplate for you and provides a simple, high-level API for both Hugging Face and ONNX backends.
pip install simpleai-sdk
from simple_ai import lille
# This will download and cache the model on first run.
# Specify the model version: "130m-instruct" (default) or "130m-base"
# Specify the backend: "huggingface" (default) or "onnx"
model = lille("huggingface", "130m-instruct")
# --- For Chat (with instruct model) ---
print("--- Chat Example ---")
response1 = model.chat("What is the capital of France?", max_new_tokens=50)
print(f"Bot: {response1}")
response2 = model.chat("And what is its population?", max_new_tokens=50, top_p=0.90)
print(f"Bot: {response2}")
# This resets the chat history
model.reset_chat()
# --- For Text Completion (with base or instruct model) ---
prompt = "Artificial Intelligence is"
response = model.generate(prompt, max_new_tokens=50, temperature=0.9)
print(f"\n--- Completion Example ---\n{prompt}{response}")
simpleai-sdk currently)You can also use the model directly with the transformers library for more advanced use cases.
pip install transformers torch simpleai-sdk
import torch
from transformers import AutoTokenizer, AutoConfig, AutoModelForCausalLM
from simple_ai.model_hf import LilleConfig, LilleForCausalLM
# 1. Register the custom model architecture with Hugging Face
AutoConfig.register("lille-130m", LilleConfig)
AutoModelForCausalLM.register(LilleConfig, LilleForCausalLM)
# 2. Define constants and setup device
MODEL = "Nikity/lille-130m-instruct"
DEVICE = "cuda" if torch.cuda.is_available() else "cpu"
# 3. Load tokenizer and model
tokenizer = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(
MODEL,
torch_dtype="auto",
device_map=DEVICE,
)
# 4. Prepare chat prompt and tokenize it
chat = [{"role": "user", "content": "What is the capital of France?"}]
inputs = tokenizer.apply_chat_template(
chat,
add_generation_prompt=True,
return_tensors="pt"
).to(DEVICE)
# 5. Generate a response
with torch.inference_mode():
outputs = model.generate(
input_ids=inputs,
max_new_tokens=512,
eos_token_id=tokenizer.eos_token_id,
pad_token_id=tokenizer.pad_token_id,
do_sample=True,
temperature=0.5,
top_p=0.95,
)
# 6. Decode and print the response
response = tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True)
print(response)
You can replicate the pretraining of Lille-130M-Base or fine-tune it on your own dataset using the provided scripts.
First, clone the repository and install the required dependencies:
git clone https://github.com/Nikityyy/lille
cd lille
pip install -r requirements.txt
Note on the Optimizer: The default Sophia-Triton optimizer requires the Triton library. Triton is officially supported on Linux with NVIDIA GPUs. While experimental installation on Windows is possible, it can be a complex and difficult process. For a much simpler setup on Windows and macOS, or if you prefer not to install Triton, it is highly recommended to use a pure PyTorch implementation of Sophia instead:
sophia_triton.py file with the code from this link.train.py script should work without any import changes, as the class name SophiaG is the same.The training script expects data in a specific .npz format containing tokenized documents and their offsets.
For Pretraining (like FineWeb-Edu):
Use the prepare_dataset_fineweb.py script. It will stream the dataset from Hugging Face, apply filters, tokenize the text, and save it in the required format.
python prepare_dataset_fineweb.py
This will create data/fineweb_edu_sample_10BT/train.npz and val.npz.
For Finetuning (Instruction Datasets):
Use the prepare_dataset.py script. Your input data should be a single .txt file where each example is separated by the <|endoftext|> token.
data/my_dataset/train.txt.input_file_path and output_dir variables in prepare_dataset.py.python prepare_dataset.py
This will create train.npz and val.npz in your specified output directory.
All training logic is handled by train.py. You can configure hyperparameters directly at the top of this file.
To Pretrain from Scratch:
train.py, set finetune = False.data_dir, batch_size, etc.python train.py
To Fine-tune a Pretrained Model:
train.py, set finetune = True.resume_checkpoint to the path of the pretrained model checkpoint (e.g., checkpoints/best_model.pt).finetune_data_dir and finetune_learning_rate.python train.py
Checkpoints will be saved in the directory specified by out_dir (for pretraining) or finetune_out_dir (for fine-tuning). The best model based on validation loss will be saved as best_model.pt.
Lille-130M-Base)sample-10BT configuration of the HuggingFaceFW/fineweb-edu dataset.Lille-130M-Instruct)Lille models primarily understand and generate content in English. While powerful for their size, they can produce text that may not always be factually accurate, logically consistent, or free from biases present in the training data. These models should be used as assistive tools rather than definitive sources of information. Users should always verify important information and critically evaluate any generated content.
Lille is a key component of my initiative to build and release a complete, truly open-source stack for language modeling. All components are designed to work together seamlessly.
This project is licensed under the Apache-2.0 License.
If you use Lille or any part of this open-source stack in your work, please consider citing it:
@misc{lille-130m,
author = {Nikita Berger},
title = {Lille: A Truly Open-Source 130M Language Model},
year = {2025},
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/Nikityyy/lille}}
}