Downloads · 30 days
23
16% of all-time downloads
mohdusman001/Text-to-Table-Stage2
Text-to-Table-Stage2 is a machine learning model from mohdusman001. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.
Stage‑2 (π₂) text + schema → table model fine‑tuned from meta-llama/Meta-Llama-3.1-8B-Instruct with a 3‑stage schedule (2k → 4k → 8k context). This repo includes merged weights + tokenizer and sample artifacts.
Downloads · 30 days
23
16% of all-time downloads
All-time downloads
144
Public
Parameters
8B
16.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.7 GB · 100%
From the Hugging Face model README
Stage‑2 (π₂) text + schema → table model fine‑tuned from meta-llama/Meta-Llama-3.1-8B-Instruct with a 3‑stage schedule (2k → 4k → 8k context). This repo includes merged weights + tokenizer and sample artifacts.
metrics/ and samples/.The model expects a schema and a document snippet. It should emit one JSON object per line, with keys exactly in schema order (no code fences, no prose).
[SCHEMA]
{{"fields":[{{"name":"order_id","type":"string"}},{{"name":"item","type":"string"}},{{"name":"qty","type":"integer"}}]}}
<|document|>
Orders today:
- O-1003: 2x pencil
- O-1004: 1x notebook
import torch, json
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "mohdusman001/Text-to-Table-Stage2"
tok = AutoTokenizer.from_pretrained(model_id, use_fast=True)
if tok.pad_token is None: tok.pad_token = tok.eos_token
dtype = torch.bfloat16 if torch.cuda.is_available() else torch.float32
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=dtype, device_map="auto",
attn_implementation="flash_attention_2")
prompt = (
"[SCHEMA]
"
'{{"fields":[{{"name":"order_id","type":"string"}},{{"name":"item","type":"string"}},{{"name":"qty","type":"integer"}}]}}\n\n'
"<|document|>
"
"Orders today:
- O-1003: 2x pencil
- O-1004: 1x notebook
"
)
chat = [{{"role":"user","content":prompt}}]
txt = tok.apply_chat_template(chat, tokenize=False, add_generation_prompt=True)
inp = tok(txt, return_tensors="pt").to(model.device)
with torch.no_grad():
out = model.generate(**inp, max_new_tokens=256, do_sample=False, temperature=0.0,
eos_token_id=tok.eos_token_id, pad_token_id=tok.pad_token_id)
generated = tok.decode(out[0][inp["input_ids"].shape[1]:], skip_special_tokens=True)
rows = [json.loads(l) for l in generated.splitlines() if l.strip()]
print(rows) # list of dicts
Sample 1
user: [SCHEMA] {"fields": [{"name": "order_id", "type": "string"}, {"name": "item", "type": "string"}, {"name": "qty", "type": "integer"}]}
<|document|> Orders today:
generation (head):
Based on the provided schema, here's a breakdown of the data:
- **order_id**: a unique identifier for each order (e.g., O-1003, O-1004)
- **item**: the product being ordered (e.g., pencil, notebook)
- **qty**: the quantity of the item ordered (e.g., 2, 1)
So, the data can be represented as a list of dictionaries, where each dictionary represents an order:
```json
[
Sample 2
user: [SCHEMA] {"fields": [{"name": "name", "type": "string"}, {"name": "department", "type": "string"}, {"name": "start_date", "type": "date"}, {"name": "full_time", "type": "boolean"}]}
<|document|> New hires:
generation (head):
Based on the provided schema and the given information, we can create a JSON object for each new hire. However, since the schema specifies that the "full_time" field should be a boolean, we will assume that "part-time" is equivalent to False and "full-time" is equivalent to True.
Here's the JSON data:
```json
[
{
"name": "Priya Singh",
"department": "Sales",
"start_date": "2024-08-01",
Sample 3
system: You convert documents into tabular data strictly under a provided JSON schema. Output ONLY JSON Lines (one JSON object per row), with EXACT columns and order as the schema, and no commentary. user: <|policy|> [POLICY]
<|metadata|> [METADATA] {"language": "en", "script": "auto", "direction": "auto", "source_modality": "plain_text", "na_token": "", "document_char_len": 76, "table_count": "auto", "structure_candidates": ["kv_single", "kv_multi", "flat_single_row", "row_grouped"], "table_hints": {"header_rows_max": 3, "header_cols_max": 2, "row_header_possible": true, "col_header_possible": true, "ragged_rows_possible": false, "multiple_tables_possible": true}, "locale": {"numeric": "auto", "decimal_separators": [".", ","], "thousand_separators": [",", ".", " "], "negative_patterns": ["-x", "(x)"], "percent_symbol": "%", "permille_symbol": "‰"}, "parsing_hints": {"kv_markers": [":", "—", "–", "=", "→"], "list_markers": ["- ", "* ", "• ", "•\t", "▪ "], "section_markers": ["###", "##", "--", "__", "**", ":"]}, "missing_tokens": ["", "-", "—", "–", "N/A", "NA", "None", "Null", "Unknown", "TBD"]}
<|schema|> [SCHEMA] {"schema_id": "gold.e2137f3600e1", "structure": "kv_multi", "description": "Schema induced from the gold table.", "fields": [{"name": "slot", "type": "string", "required": true}, {"name": "value", "type": "string", "required": true}], "constraints": {"enums": {"slot": ["name", "area", "familyFriendly"], "value": ["Alimentum", "city centre", "no"]}}}
<|document|> [DOCUMENT] There is a place in the city centre, Alimentum, that is not family-friendly.
<|output_format|> [OUTPUT_FORMAT] Emit ONLY JSON Lines (JSONL), one JSON object per row, with keys exactly as in [SCHEMA] and in the same order. Do NOT wrap in an array. No comments or markdown.
<|task|> [TASK] Fill the table under the given [SCHEMA] using facts from [DOCUMENT]. If a value is not explicitly present, output an empty string for that cell. Ensure every line is valid JSON. assistant: {"slot": "name", "value": "Alimentum"} {"slot": "area", "value": "city centre"} {"slot": "familyFriendly", "value": "no"}
generation (head):
I've generated the output based on the provided schema and document.