Downloads · 30 days
0
jacob-valdez/tensorcode-chatbot-cognitive-experimental-001
tensorcode-chatbot-cognitive-experimental-001 is a machine learning model from jacob-valdez. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for tensorcode.
This revision replaces the earlier weights on main. The earlier weights were trained before the bounded workspace update (memoryupdate="relativermsbounded") and load only with TensorCode source commit 6607a8b. Revisio…
Downloads · 30 days
0
Access
Public
Updated Sep 24, 2026
Parameters
517M
4.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.1 GB · 99%
How the weights are stored.
F32517M · 100%
From the Hugging Face model README
This revision replaces the earlier weights on main. The earlier weights were
trained before the bounded workspace update (memory_update="relative_rms_bounded")
and load only with TensorCode source commit 6607a8b. Revision
8836ba59275dc6d8ceeb04462b4191beb9813452 remains available, unchanged, for that
runtime, together with its diagnostic files (diagnostic-raw.json,
hub-first-attempt-failure.json), which describe the earlier weights and are not
carried forward on main.
| Revision | Architecture | Loads with |
|---|---|---|
main (this card; bounded retrain published 2026-09-24) | bounded workspace update | TensorCode 0.4.0a4 (source commit 090ebc4; the PyPI wheel's package files are identical) |
8836ba59275dc6d8ceeb04462b4191beb9813452 (earlier) | earlier unbounded update | TensorCode commit 6607a8b only |
This checkpoint owns its language encoder/decoder, hypothesis generator, source verifier, ranking operations, workspace and retrieval encoder. It does not call a hosted inference provider. The caller supplies evidence explicitly. Generated statements remain hypotheses, and the model may abstain.
Two components were retrained on the current bounded-workspace architecture:
The other components are reused byte-for-byte:
jacob-valdez/tensorcode-investigator-hotpot-001);from tensorcode.tools.chatbot import Chatbot
bot = Chatbot.from_pretrained("jacob-valdez/tensorcode-chatbot-cognitive-experimental-001")
session = bot.new_session()
answer = session({
"question": "Your question",
"evidence": [{"id": "source-1", "source_id": "document-name", "text": "Actual source text"}],
})
print(answer)
print(session.last_result)
Pin the commit hash from this repository's history for reproducible loading. Use independent sessions for independent evidence histories. Saved model weights exclude sessions and conversations. Training and optimizer checkpoints are separate artifacts.
The 32 questions are reused. They are the same 32 HotpotQA
distractor-validation questions (rows 272–303, revision 1908d6af) with oracle
supporting passages that the previous revision was evaluated on. Later development
work also reused them. This is a fixed-configuration re-run on known cases, not a
held-out test. Nothing was tuned on these cases. Configuration, component and
source hashes were frozen before the run (final-freeze.json).
| 32 fixed questions | previous revision (8836ba5) | this revision |
|---|---|---|
| Answered (not abstained) | 2 | 3 |
| Source-reviewed correct | 1 | 2 |
| Incorrect or non-answer | 1 circular non-answer | 1 incorrect answer |
| Abstentions | 30 (93.75%) | 29 (90.625%) |
| Selected statement preserved in every answer | yes | yes (3/3) |
| Failed calls, truncations, veto violations, uncalibrated records | 0 | 0 |
Source-grounded review of the three answers (assistant review, not independent
human annotation; see manual-factual-review.json):
So 2/32 answers are correct overall. Three answered cases are too few to estimate reliability. The one error is a wrong factual claim, a worse kind of error than the previous revision's circular non-answer.
Controls and retrieval:
Repository-document smoke (docs/pretrained.md): the source was retrieved again
across a new episode and after session save/load. The model did not abstain. It
answered "The preferred model host for TensorCode is nugging Face .", which corrupts
"Hugging Face", and the verifier accepted it (0.864 support). The previous revision
abstained on its smoke, but the document excerpt had changed between the runs
(sha256 a345be80… now, 0f2a4feb… then). Given the earlier excerpt, this
revision also abstained, including after new_episode and session save/load
(smoke-previous-excerpt.json). Document QA competence is not established.
Component evaluations use the same splits as the previous revision and the
current versions of the scripts (examples/train_hypotheses.py sha256 a5ce6854…,
previously 526b6046…; proposal module src/tensorcode/_internal/proposals.py fc8ab24f…, previously
6b2e2e31…).
Previous scores are in parentheses.
Other pipeline metrics (previous in parentheses): whole-word answer containment 0.0625 (0.03125, a diagnostic, not accuracy), candidate substring coverage 0.594 (0.594), realization NLI support on the selected statement 0.874 (0.827), format-sensitive short-answer exact match 0.0 (0.0).
The complete model roundtrips through save_pretrained/from_pretrained with a
bitwise-equal state dict and identical configuration bytes (the re-saved
safetensors header can order tied-weight aliases differently). In a fresh process,
the complete cognitive receipts of both correct cases reproduced exactly, including
session save/load and new_episode.
All runs used unmodified TensorCode source at commit 090ebc4 (0.4.0a4) on an
NVIDIA GB10 with torch 2.14.0+cu130 and transformers 5.17.0. Exact commands are in
run-manifest.json; every training report is in training-reports/.
Generator. Script: examples/train_hypotheses.py.
turker_answer declarations joined to
the original SQuAD paragraphs, split by article (same split hashes as before).How the generator was chosen (full disclosure):
generator-selection-v2.json, generator-extended-validation.json).training-reports/lr5e-5-replicates/). So lr 5e-5 cleared the bar in 3 of 3
seeds; this does not undo the post hoc choice of learning rate.Realizer. Script: examples/train_realization.py.
Apart from memory_update: relative_rms_bounded (top level and generator), the
configuration records two fields the previous configuration did not have:
cognition.conversation_context_tokens: 128 and
cognition.investigator.verification_scope: "source" (current-source defaults).
Policy, memory, proposal count and foundation revisions are unchanged. Foundation
repository strings record the relative local paths used during assembly (for
example ../artifacts/minilm-foundation); loading does not use them.
This model does not establish general cognition, reliable multi-hop reasoning or factual guarantees.
Foundation licenses and dataset terms still apply. SQuAD is CC-BY-SA-4.0, and the QA2D mirror declares MIT.
verifier-snli.json); sha256 2a97c54e…Files: final-freeze.json, assembly-provenance.json, evaluation.json (the raw
report), qualification.json (gate-by-gate comparison with the previous
revision), manual-factual-review.json, fresh-process-replay.json,
data-manifest.json, smoke-previous-excerpt.json, hypotheses-qa2d.json (selected generator),
realization-qa2d.json, verifier-snli.json, run-manifest.json and
training-reports/. No parameters or policy thresholds changed after the final
evaluation began.