Downloads · 30 days
31
100% of all-time downloads
ItsnotAilabs/HAM-384
HAM-384 is a sentence similarity model from ItsnotAilabs. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as apache-2.0.
Downloads · 30 days
31
100% of all-time downloads
All-time downloads
31
Public
Parameters
26.7M
196 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors107 MB · 97%
From the Hugging Face model README
HAM-384 is a highly compact embedding model designed for rapid memory retrieval, short-term to long-term memory (STM/LTM) consolidation, and enabling Hebbian spreading activation cascades in intelligent systems.
Fine-tuned from BAAI/bge-small-en-v1.5, HAM-384 is optimized for speed and associative recall. Its primary function is to act as the core vector retrieval mechanism in memory cascades, where memory nodes activate related nodes based on temporal and semantic proximity.
| Format | Precision | RAM Required | Latency (ms/query) |
|---|---|---|---|
| PyTorch | FP32 (Full) | ~130 MB | ~4.1 |
| PyTorch | FP16 | ~65 MB | ~2.5 |
| ONNX | INT8 | ~33 MB | ~1.3 |
| GGUF | Q4_K_M | ~20 MB | ~0.8 |
The model was fine-tuned on associative memory traces:
Optimized for associative recall and resonance matching rather than strict semantic equivalence, enabling the system to jump between temporally and contextually linked concepts.
HAM-384 sacrifices some general retrieval performance for highly specialized associative recall:
| Benchmark | Metric | Score |
|---|---|---|
| MTEB Retrieval | Average | 51.8 |
| Memory-Recall@3 (Custom) | Accuracy | 89.4% |
| Activation-Cascade (Custom) | F1-Score | 82.1% |
| Peak Throughput | Docs/sec | ~74,222 |
Note: Memory-Recall@3 and Activation-Cascade F1 are custom metrics specific to internal MedinaMemorySystems evaluations.
For queries mimicking memory retrieval, no special prefix is enforced, though matching the style of associative logs helps recall:
{memory_fragment}
from sentence_transformers import SentenceTransformer, util
# Load the model
model = SentenceTransformer('MedinaMemorySystems/ham-384')
# A target memory to recall from
current_memory = "The user configured the security firewall rules for the cloud database."
# Memory bank (LTM)
memory_bank = [
"Database backup completed successfully at midnight.",
"Firewall updated to block unauthorized external IP ranges.",
"User logged in from a new device."
]
# Encode the current state and memory bank
query_embedding = model.encode(current_memory)
corpus_embeddings = model.encode(memory_bank)
# Find top-3 nearest neighbors for spreading activation
hits = util.semantic_search(query_embedding, corpus_embeddings, top_k=3)[0]
print("Top Associative Retrievals:")
for hit in hits:
print(f"- {memory_bank[hit['corpus_id']]} (Score: {hit['score']:.4f})")
import torch
import torch.nn.functional as F
from transformers import AutoTokenizer, AutoModel
# Load model and tokenizer
tokenizer = AutoTokenizer.from_pretrained('MedinaMemorySystems/ham-384')
model = AutoModel.from_pretrained('MedinaMemorySystems/ham-384')
# Tokenize inputs
sentences = ["The user configured the security firewall rules."]
encoded_input = tokenizer(sentences, padding=True, truncation=True, return_tensors='pt')
# Compute token embeddings
with torch.no_grad():
model_output = model(**encoded_input)
# Perform mean pooling
def mean_pooling(model_output, attention_mask):
token_embeddings = model_output[0]
input_mask_expanded = attention_mask.unsqueeze(-1).expand(token_embeddings.size()).float()
return torch.sum(token_embeddings * input_mask_expanded, 1) / torch.clamp(input_mask_expanded.sum(1), min=1e-9)
sentence_embeddings = mean_pooling(model_output, encoded_input['attention_mask'])
sentence_embeddings = F.normalize(sentence_embeddings, p=2, dim=1)
print(sentence_embeddings)
@misc{medinamemorysystems2026ham,
title={HAM-384: A Compact BERT Model for Hybrid Associative Memory Cascades},
author={MedinaMemorySystems},
year={2026},
publisher={Hugging Face}
}