Downloads · 30 days
429
4% of all-time downloads
pramodkoujalagi/SmolLM2-360M-Instruct-Text-2-JSON
SmolLM2-360M-Instruct-Text-2-JSON is a text generation model from pramodkoujalagi. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
429
4% of all-time downloads
All-time downloads
10.1K
Public
Repo size
2.2 GB
Likes
0
Public
Click a slice to open those files.
.gguf2.2 GB · 51%
From the Hugging Face model README
Developed by: Pramod Koujalagi
A fine-tuned version of SmolLM2-360M-Instruct-bnb-4bit specialized for parsing unstructured calendar event requests into structured JSON data.
This model is fine-tuned on SmolLM2-360M-Instruct-bnb-4bit using QLoRA to extract structured calendar event information from natural language text. It identifies and structures key scheduling entities like action, date, time, attendees, location, duration, recurrence, and notes.
You can use the SmolLM2-360M-Instruct-Text-2-JSON model to parse natural language event descriptions into structured JSON format.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
import json
# Load model and tokenizer
model_name = "pramodkoujalagi/SmolLM2-360M-Instruct-Text-2-JSON"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name)
def parse_calendar_event(text):
# Format the prompt
formatted_prompt = f"""<|im_start|>user
Extract the relevant event information from this text and organize it into a JSON structure with fields for action, date, time, attendees, location, duration, recurrence, and notes. If a field is not present, return null for that field.
Text: {text}
<|im_end|>
<|im_start|>assistant
"""
# Generate response
inputs = tokenizer(formatted_prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=512,
do_sample=True,
temperature=0.1,
top_p=0.95,
pad_token_id=tokenizer.eos_token_id
)
# Process response
output_text = tokenizer.decode(outputs[0], skip_special_tokens=False)
response = output_text.split("<|im_start|>assistant\n")[1].split("<|im_end|>")[0].strip()
# Return formatted JSON
parsed_json = json.loads(response)
return json.dumps(parsed_json, indent=2)
# Example input
event_text = "Plan an exhibition walkthrough on 15th, April 2028 at 3 PM with Harper, Grace, and Alex in the art gallery for 1 hour, bring bag."
# Output
print("Prompt:")
print(event_text)
print("\nModel Output:")
print(parse_calendar_event(event_text))
Output
Prompt:
Plan an exhibition walkthrough on 15th, April 2028 at 3 PM with Harper, Grace, and Alex in the art gallery for 1 hour, bring bag.
Model Output:
{
"action": "Plan an exhibition walkthrough",
"date": "15/04/2028",
"time": "3:00 PM",
"attendees": [
"Harper",
"Grace",
"Alex"
],
"location": "art gallery",
"duration": "1 hour",
"recurrence": null,
"notes": "Bring bag"
}
<!--
**Example:**
```
Input: "Plan an exhibition walkthrough on 15th, April 2028 at 3 PM with Harper, Grace, and Alex in the art gallery for 1 hour, bring bag."
```
```
Output: {
"action": "Plan an exhibition walkthrough",
"date": "15/04/2028",
"time": "3:00 PM",
"attendees": [
"Harper",
"Grace",
"Alex"
],
"location": "art gallery",
"duration": "1 hour",
"recurrence": null,
"notes": Bring bag
}
``` -->
The model was trained on a custom dataset consisting of 1,149 examples (1,034 training, 115 validation) of natural language event descriptions paired with structured JSON outputs. The dataset includes a wide variety of event types, date/time formats, and varying combinations of fields.