Downloads · 30 days
11
18% of all-time downloads
NebuIA/nebuia_extract_small
nebuia_extract_small is a text generation model from NebuIA. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
<p align="center" <img src="Brain.png" width="600" </p
Downloads · 30 days
11
18% of all-time downloads
All-time downloads
62
Public
Parameters
1.5B
3.1 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors3.1 GB · 100%
From the Hugging Face model README
nebuia_extract_small is an extraction model inspired by NuExtract. nebuia_extract_small is a version of qween 1.5b, fine-tuned on a private high-quality synthetic dataset for entity extraction in Spanish legal texts with an 8k context length. Supports JSON template like nu extract describing the information you need to extract. NebuIA Extract specializes in identifying and extracting legal entities and relevant information from Spanish legal documents.
Same template as NuExtract
import json
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
def predict_extract(model, tokenizer, text, schema):
schema = json.dumps(json.loads(schema), indent=4)
input_llm = "<|input|>\n### Template:\n" + schema + "\n"
input_llm += "### Text:\n"+text +"\n<|output|>\n"
input_ids = tokenizer(input_llm, return_tensors="pt", truncation=True, max_length=4000).to("cuda")
output = tokenizer.decode(model.generate(**input_ids)[0], skip_special_tokens=True)
return output.split("<|output|>")[1].split("<|end-output|>")[0]
model = AutoModelForCausalLM.from_pretrained("NebuIA/nebuia_extract_small", trust_remote_code=True, torch_dtype=torch.bfloat16)
tokenizer = AutoTokenizer.from_pretrained("NebuIA/nebuia_extract_small", trust_remote_code=True)
model.to("cuda")
model.eval()
text = """large legal text"""
schema = """{
"calusulas": [],
"notario": "",
"jurisdiccion": {
"clausula_jurisdiccion": "",
"lugar": ""
}
}"""
prediction = predict_extract(model, tokenizer, text, schema)
print(prediction)