Downloads · 30 days
0
jumafernandez/trace-checkpoints
trace-checkpoints is a machine learning model from jumafernandez. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as mit.
Checkpoints for TRACE, a contextual turn encoder that learns the update rule applied to frozen turn embeddings in task-oriented dialogue.
Downloads · 30 days
0
Access
Public
Updated Aug 12, 2026
Repo size
15.7 GB
Likes
0
Public
Click a slice to open those files.
.safetensors17.8 GB · 100%
From the Hugging Face model README
Checkpoints for TRACE, a contextual turn encoder that learns the update rule applied to frozen turn embeddings in task-oriented dialogue.
A frozen base encoder $f_1$ maps each utterance to a static vector $e_t$; TRACE ($f_2$) is a Transformer whose tokens are turns, and it maps the sequence $(e_1,\dots,e_T)$ to contextual representations $h_t$. Training is self-supervised and requires no functional annotation.
All checkpoints come from the REPRO campaign: a pre-specified protocol in which dialogues are partitioned by near-duplicate cluster, probes are fit on a held-out development split, and test labels never inform training, checkpoint or hyperparameter selection. Every model was retrained from scratch under that partition.
| Canonical model | trace-repro-lite-ar-s42/best |
| Base encoder (frozen) | sergioburdisso/dialog2flow-joint-bert-base |
| Attention mode | autoregressive — the deployable one |
| Parameters | 44.9M (6 layers, 8 heads, 768-dim) |
The AR mode is the one to use in practice: it only attends to the dialogue so far. Bidirectional checkpoints see the whole window and are included for the representation-vs-anticipation analysis.
trace-repro-<recipe>-<mode>-s<seed>[-xlc][-scratch]
│ │ │ │ └── trained from scratch on the 28M curriculum
│ │ │ └───────── continued pretraining on ~28M turns
│ │ └──────────────── training seed: 42, 7, 123
│ └─────────────────────── ar | bidi
└──────────────────────────────── lite (6 layers) | deep (12) | gru (recurrent control)
Base-ablation checkpoints carry the base in the name instead: mpnet, todbert.
| checkpoints | what they support |
|---|---|
trace-repro-{lite,deep}-{ar,bidi}-s{42,7,123} | main ladder: current- and next-act prediction |
trace-repro-gru-ar-s{42,7,123} | learned recurrent control — isolates attention from recurrence |
trace-repro-{mpnet,todbert}-lite-ar-s{42,7,123} | three-base ablation |
trace-repro-lite-ar-s*-xlc, -xlc-scratch | data-scale study (28M-turn curriculum) |
trace-repro-deep-ar-s*-xlc | model + data scale combined |
Untrained controls are not published: they are randomly initialized copies of the same architectures, reproducible from the configs with seeds 0–4.
<checkpoint>/
├── best/ # selected by validation loss — use this one
│ ├── config.json
│ └── model.safetensors
├── config.json # final epoch
├── model.safetensors
└── trainlog.jsonl # per-epoch train/val loss
from contextual_turn_embeddings import ContextualTurnModelV2, encode_dialogues
model = ContextualTurnModelV2.from_pretrained("best", device="cpu").eval()
# embeddings: (n_turns, 768) from the frozen base, in dialogue order
H, meta = encode_dialogues(model, frames, embeddings=E, device="cpu")
The package lives at
packages/contextual-turn-embeddings
in the project monorepo, together with the partition, the training and evaluation
scripts, the seeds and the commands that regenerate every table.
The precomputed $e_t$ vectors used as input are published separately as datasets:
jumafernandez/d2f-turn-embeddings-*.
Fernández, J. M., Errecalde, M., & Burdisso, S. TRACE: Learning the Update for Contextual Turn Representations in Task-Oriented Dialogue. (under review)