Downloads · 30 days
8
7% of all-time downloads
Lux-In-Tenebris/research_summary_deconstructor
research_summary_deconstructor is a machine learning model from Lux-In-Tenebris. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Fine-tuned version of IBM Granite 3.3 2b base
Downloads · 30 days
8
7% of all-time downloads
All-time downloads
123
Public
Parameters
2.5B
5.1 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors5.1 GB · 100%
From the Hugging Face model README
Fine-tuned version of IBM Granite 3.3 2b base
Accepts a research project title and summary and creates a largely extractive breakdown (with some minor tweaks for coherence) into:
with seperate keywords for each of these breakdowns.
Additionally includes lists of:
Usage:
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
import json
input_title: str = "your project title here"
input_summary: str = "your summary here"
model_name: str = "Lux-In-Tenebris/research_summary_deconstructor"
device = "cuda" if torch.cuda.is_available() else "cpu"
model = AutoModelForCausalLM.from_pretrained(model_name, dtype=torch.bfloat16).to(device)
tokenizer = AutoTokenizer.from_pretrained(model_name)
model.eval()
input_text = json.dumps({"title": input_title, "summary": input_summary}, ensure_ascii=False)
inputs = tokenizer(input_text, return_tensors="pt", truncation=False, padding=False).to(device)
with torch.no_grad():
outputs = model.generate(**inputs, max_new_tokens=8_192, temperature=1.0, do_sample=True)
generated_tokens = outputs[0][inputs["input_ids"].shape[1]:]
decoded = tokenizer.decode(generated_tokens, skip_special_tokens=True)
Will of course need adjustment for batch processing, and highly recommend checking for invalid json and additional characters being output.
Trained on approx 5000 synthetic research project titles, summaries, and structured elements, for 2 epochs using an RTX6000 Pro