Downloads · 30 days
0
sahil239/falcon-lora-chatbot
falcon-lora-chatbot is a machine learning model from sahil239. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Model Description This is a LoRA adapter for the Falcon architecture, fine-tuned on domain-specific chat-style data for enhanced language understanding and generation. It was built using the PEFT library with 4-bit qu…
Downloads · 30 days
0
Access
Public
Updated Aug 7, 2025
Repo size
6.3 MB
Likes
0
Public
Click a slice to open those files.
.safetensors6.3 MB · 57%
From the Hugging Face model README
Model Description
This is a LoRA adapter for the Falcon architecture, fine-tuned on domain-specific chat-style data for enhanced language understanding and generation. It was built using the PEFT library with 4-bit quantization.
Direct Use
This adapter is intended to be used with Falcon base models to improve instruction-following and chatbot-like behavior on English-language prompts. It is suitable for:
Downstream Use
Can be further fine-tuned for more specific domains such as finance, DIY assistance, or medical Q&A, depending on your dataset.
Out-of-Scope Use
Not suitable for real-time critical decision-making tasks such as:
As with all large language models, outputs may reflect biases in the training data. The adapter may reproduce toxic, biased, or incorrect information and should be monitored in production use.
Users should:
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel, PeftConfig
# Load base Falcon model and tokenizer
base_model = AutoModelForCausalLM.from_pretrained("tiiuae/falcon-7b", device_map="auto", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("tiiuae/falcon-7b")
# Load LoRA adapter
adapter = PeftModel.from_pretrained(base_model, "sahildesai/falcon-lora")
# Run inference
inputs = tokenizer("Explain black holes to a 12-year-old.", return_tensors="pt").to("cuda")
outputs = adapter.generate(**inputs, max_new_tokens=100)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Training Data
The model was fine-tuned on a subset of conversational and instruction-following datasets derived from public chat data.
Preprocessing
Training Hyperparameters
query_key_valueSpeeds, Sizes, Times
adapter_model.bin)Testing Data
Subset of instruction-following prompts held out during training.
Factors
Evaluation included:
Metrics
Results
Model Examination (optional)
A sample comparison between the base and adapter model showed that the adapter improved clarity and tone in responses.
Hardware Type: NVIDIA A100 40GB (Google Colab Pro)
Hours used: ~2.5 hours
Cloud Provider: Google
Compute Region: US
Carbon Emitted: ~2.1 kg CO₂ (estimated via ML CO₂ calculator)
Model Architecture and Objective
Compute Infrastructure
Hardware
Software
transformers==4.38.2peft==0.16.0accelerate, datasets, bitsandbytesBibTeX:
@misc{desai2025falconlora,
title={Falcon LoRA Adapter},
author={Sahil Desai},
year={2025},
url={https://huggingface.co/sahildesai/falcon-lora}
}
Model Card Authors: Sahil Desai
Model Card Contact: https://sahildesai.dev / [Hugging Face profile]