Downloads · 30 days
0
SentiChain/aparecium-v2-pooled-reverser
aparecium-v2-pooled-reverser is a text generation model from SentiChain. Use it when you need the model to write or continue text. It is set up for pytorch. The card lists the license as mit.
- Task: Reconstruct natural-language crypto social-media posts from a single pooled MPNet embedding (reverse embedding). - Focus: Crypto domain (social-media posts / short-form content). - Checkpoint: apareciumv2s1.pt…
Downloads · 30 days
0
Access
Public
Updated Nov 17, 2025
Repo size
1.1 GB
Likes
0
Public
Click a slice to open those files.
.pt1.1 GB · 100%
From the Hugging Face model README
aparecium_v2_s1.pt — S1 supervised baseline, trained on synthetic crypto social-media posts.all-mpnet-base-v2 vector of shape (768,), not a token-level (seq_len, 768) matrix.SentiChain/aparecium-seq2seq-reverser).This is a pooled-embedding variant of Aparecium, distinct from the original token-level seq2seq reverser described in
SentiChain/aparecium-seq2seq-reverser.
Reconstruction quality depends heavily on:
sentence-transformers/all-mpnet-base-v2),On the encoder side, we assume a pooled MPNet encoder:
sentence-transformers/all-mpnet-base-v2 (768‑D pooled output).On the decoder side, v2 uses the Aparecium components:
e ∈ R^768.H ∈ R^{B × S × D} suitable for a transformer decoder (multi‑scale).e.d_model = 768n_layer = 12n_head = 8d_ff = 3072H as cross‑attention memory and generates text tokens.Decoding:
r(x, e) for reranking candidates.The aparecium_v2_s1.pt checkpoint contains the adapter, sketcher, decoder, and tokenizer name, matching the training repo layout.
tweets.db).sentence-transformers/all-mpnet-base-v2:
embedding ∈ R^768 (pooled), L2‑normalized.No real social‑media content is used; all posts are synthetic, similar in spirit to the v1 project
SentiChain/aparecium-seq2seq-reverser.
This checkpoint corresponds to S1 supervised training only (no SCST/RL):
all-mpnet-base-v2 and normalized.aparecium_v2_s1.pt once training plateaus on validation cross‑entropy.Future work (not in this checkpoint):
r, repetition penalty, and entity coverage.This repo does not include a full eval harness. The S1 baseline was validated qualitatively:
all-mpnet-base-v2 (pooled, normalized).For a v1‑style, large‑scale evaluation (crypto/equities split, cosine statistics, degeneracy rate, domain drift), refer to the v1 model card:
SentiChain/aparecium-seq2seq-reverser.
Input (v2, S1 baseline):
(768,), L2‑normalized.sentence-transformers/all-mpnet-base-v2 from sentence-transformers.Do not pass a token‑level (seq_len, 768) matrix – that is the contract for the v1 seq2seq model
SentiChain/aparecium-seq2seq-reverser, not this checkpoint.
Usage pattern (high level, pseudocode):
import torch, json
from sentence_transformers import SentenceTransformer
# 1) Pooled MPNet embedding
mpnet = SentenceTransformer("sentence-transformers/all-mpnet-base-v2",
device="cuda" if torch.cuda.is_available() else "cpu")
text = "Ethereum L2 blob fees spiked after EIP-4844; MEV still shapes order flow."
e = mpnet.encode([text], convert_to_numpy=True, normalize_embeddings=True)[0] # (768,)
# 2) Load Aparecium v2 S1 checkpoint
ckpt = torch.load("aparecium_v2_s1.pt", map_location="cpu")
# 3) Recreate models from the Aparecium codebase (not included in this HF repo)
# from aparecium.aparecium.models.emb_adapter import EmbAdapter
# from aparecium.aparecium.models.decoder import RealizerDecoder
# from aparecium.aparecium.models.sketcher import Sketcher
# from aparecium.aparecium.utils.tokens import build_tokenizer
# and run the same decoding logic as in `aparecium/infer/service.py` or
# `aparecium/scripts/invert_once.py`.
# 4) Use beam search / constraints / reranking as in the training repo.
To actually use the model, you need the Aparecium codebase (training repo) where the EmbAdapter, Sketcher, RealizerDecoder, constraints, and decoding functions are defined.
To reproduce or extend this checkpoint:
train/val/test JSONL.all-mpnet-base-v2 (pooled 768‑D) and save as JSONL with {"text","embedding","plan"} fields.batch_size ≈ 64, max_len ≈ 96, lr ≈ 3e-4, cosine scheduler, warmup steps.r for reranking.If you use this model or the Aparecium codebase, please cite:
Aparecium v2: Pooled MPNet Embedding Reversal for Crypto Tweets
SentiChain (Aparecium project)
You may also reference the v1 baseline model card:
SentiChain/aparecium-seq2seq-reverser.