Downloads · 30 days
8
8% of all-time downloads
andthattoo/subquery-SmolLM
subquery-SmolLM is a text generation model from andthattoo. Use it when you need the model to write or continue text. The card lists the license as mit.
This dude is trained for boosting the performance of your RAG based question answering app
Downloads · 30 days
8
8% of all-time downloads
All-time downloads
101
Public
Parameters
152M
304 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors304 MB · 99%
From the Hugging Face model README
This dude is trained for boosting the performance of your RAG based question answering app
My motivation was to tackle a core problem of RAG with an extremely lightweight, but capable model.
If queries are
Training data was generated with Dria: A decentralized p2p network for synthetic data. Join discord to help decentralized data generation.
Heads up: Ollama version works 160 tps on 1 CPU core. No GPU? No worries. This little dude’s got you.
Use the model:
from transformers import AutoModel, AutoConfig, AutoTokenizer, AutoModelForCausalLM
config = AutoConfig.from_pretrained("andthattoo/subquery-SmolLM")
tokenizer = AutoTokenizer.from_pretrained("andthattoo/subquery-SmolLM")
model = AutoModelForCausalLM.from_pretrained("andthattoo/subquery-SmolLM", torch_dtype=torch.bfloat16)
if tokenizer.pad_token is None:
tokenizer.pad_token = tokenizer.eos_token
model.config.pad_token_id = model.config.eos_token_id
input_data = "Generate subqueries for a given question. <question>What is this?</question>"
inputs = tokenizer(input_data, return_tensors='pt')
output = model.generate(**inputs, max_new_tokens=100)
decoded_output = tokenizer.decode(output[0], skip_special_tokens=True)
Also created a python package for ease of use
pip install subquery
from subquery import TransformersSubqueryGenerator
# Using the Transformers backend
generator = TransformersSubqueryGenerator()
result = generator.generate("What is this?")
print("Follow-up questions:", result.follow_up)
print("Subqueries:", result.subquery)
or
from subquery import OllamaSubqueryGenerator
# Using the Ollama backend
generator = OllamaSubqueryGenerator()
result = generator.generate("Are the Indiana Harbor and Ship Canal and the Folsom South Canal in the same state?")
print("Follow-up questions:", result.follow_up)
print("Subqueries:", result.subquery)