Downloads · 30 days
4.2K
49% of all-time downloads
HamoAI/hamo-score-0.6b
hamo-score-0.6b is a text generation model from HamoAI. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as other.
A 0.6B model that reads one message from a wellness conversation, in Chinese or English, and returns five scores as one JSON object: Agency, Withdrawal, Extremity, Hostility, Boundary (AWEHB), each 0.0–3.0. It does no…
Downloads · 30 days
4.2K
49% of all-time downloads
All-time downloads
8.5K
Public
Parameters
596M
9.2 GB on disk
Likes
2
Public
Click a slice to open those files.
.gguf2.6 GB · 68%
From the Hugging Face model README
A 0.6B model that reads one message from a wellness conversation, in Chinese or English, and returns five scores as one JSON object: Agency, Withdrawal, Extremity, Hostility, Boundary (AWEHB), each 0.0–3.0. It does not write replies; its scores feed deterministic code (stress update, state buckets, action gating). 中文说明见下方。
pip install hamo-score, Apache-2.0): prompt format, upstream keyword gate, one-command Docker server.Message: 今天试着出门散了个步 (tried going out for a walk today), after the assistant asked 这周过得怎么样? (how was your week?)
Scores (v10):
{"A": 1.0, "W": 0.0, "E": 0.0, "H": 0.0, "B": 1.5}
A one-off walk is mid-band Agency; Boundary also credits a report of positive action.
pip install hamo-score
score_message() runs the keyword gate first and skips scoring if it fires.docker compose up in server/ of the toolkit repository serves v10 as POST /score on port 8080.python eval/run_exam.py in a clone of the repository, against your own ollama endpoint (--base-url); the Docker server exposes only port 8080.gguf/hamo-score-0.6b-v10.q8.gguf. Sampling must be neutral (temperature 0, repeat_penalty 1.0, top_k 0, top_p 1.0): ollama's default repeat_penalty 1.1 pushes scores away from zero. The toolkit's OllamaClient sends the last three with each request.Prompt template, transformers snippet and Modelfile: technical record.
| Score | What it reads (0.0–3.0 in steps of 0.5) | |
|---|---|---|
| A | Agency | Moving one's situation, or the helping relationship, forward: a decision, asking for help, a sustained healthy habit |
| W | Withdrawal | Giving up, avoiding, disengaging |
| E | Extremity | Catastrophizing chains, all-or-nothing thinking |
| H | Hostility | Attacking someone; venting without a target is not hostility |
| B | Boundary | Speaking clearly from an "I" position: a need, a limit, a stated position; calm self-description and reports of positive action also score high |
The scores are not meant to be read raw.
| Measure | Result (v10, shipped q8 GGUF, llama.cpp with Metal, temperature 0) | Measured on |
|---|---|---|
| Same state bucket as the reference labels | 435 of 453 turns (96.0%) | real final exam |
| W, E and H within ±0.5 of the reference labels | 1,204 of 1,359 judgments | real final exam |
| A: the higher-agency message of a pair scores higher | 149 of 154 pairs | synthetic A exam |
| B: the higher-boundary message of a pair scores higher | 179 of 180 pairs | synthetic B exams |
| Scores within ±0.5 of the label, mean of the five | 87.3% | toolkit self-check, 195 synthetic questions; certifies wiring, not model quality |
| Unparseable outputs | 0 | six exams |
The real final exam is 453 pseudonymised turns from three consenting internal staff members, so these are not accuracy figures for the general population. The synthetic exams are in-distribution: passing shows no regression inside the known range, not generalisation.
Not a chatbot, not a diagnostic instrument, not a crisis detector. It does not detect or handle crises, and none of its scores, W included, is a crisis signal. Crisis handling is the job of deterministic code that runs before the model: the toolkit's keyword gate is a starting point, not a complete screen, and LICENSE §3c makes an independent upstream mechanism a condition of consumer-facing mental-wellness deployments.
v10 met 17 of the 18 checks of its signed pre-registration. The one it did not meet, G5a, was a stand-in for crisis handling, which the founder ruled outside this model after seeing the result; under the registration as signed the verdict was "rejected", and v10 is released by the founder's decision. The reported checks, exams and method are in the technical record.
Synthetic conversations labelled by an open-weight teacher model, plus 440 pseudonymised turns from consenting internal staff. External user conversations never enter training. Training messages: 76.0% Chinese, 13.1% English, 10.9% mixed; results were not broken down by language.
main since 2026-10-06 (UTC). Pin it with revision="f3869312f0222992c3eb2e938e78090de38eb80a".gguf/hamo-score-0.6b-v10.q8.gguf; the v9, v7 and v6.1 GGUFs stay in gguf/. Digests and rollback commands: technical record.给每句话把脉的小模型:0.6B,读心理支持对话里的一条消息(中文、英文都行),输出一个 JSON 对象,里面是五个分数:行动力、退缩、极端化、敌意、边界感(AWEHB),各 0.0–3.0。它只打分,不写回复;分数是给确定性代码用的(压力更新、状态桶、动作门控)。
pip install hamo-score,Apache-2.0):提示词格式、放在模型上游的关键词闸门、一条命令启动的 Docker 服务。消息: 今天试着出门散了个步(上一句是助手问「这周过得怎么样?」)
打分(v10):
{"A": 1.0, "W": 0.0, "E": 0.0, "H": 0.0, "B": 1.5}
单次散步算中档的行动力;边界感的口径也给「报告积极行动」记分。
pip install hamo-score
score_message() 先跑关键词闸门,命中就不评分。server/ 下执行 docker compose up,8080 端口 POST /score,跑的是 v10。python eval/run_exam.py,对着自己的 ollama 服务(--base-url);Docker 服务只开放 8080 端口。gguf/hamo-score-0.6b-v10.q8.gguf。采样必须取中性值(温度 0、repeat_penalty 1.0、top_k 0、top_p 1.0):ollama 默认的 repeat_penalty 1.1 会把分数从 0 往上推。工具包的 OllamaClient 每次请求都带后三项。提示词模板、transformers 代码和 Modelfile 见技术档案。
| 分数 | 读的是什么(0.0–3.0,步长 0.5) | |
|---|---|---|
| A | 行动力 | 主动推动处境或疗愈关系向前:做决定、主动求助、坚持健康习惯 |
| W | 退缩 | 放弃、回避、抽离 |
| E | 极端化 | 灾难化连锁、非黑即白 |
| H | 敌意 | 攻击某个人;没有对象的抱怨不算 |
| B | 边界感 | 站在「我」的位置把需要、界限、立场说清楚;平静有条理的自述、报告积极行动也给高分 |
分数不是直接读的。
| 指标 | 结果(v10,随包 q8 GGUF,llama.cpp + Metal,温度 0) | 考卷 |
|---|---|---|
| 与参照标签落在同一个状态桶 | 453 轮中 435 轮(96.0%) | 真实终评 |
| W、E、H 与参照标签相差不超过 0.5 | 1,359 个判断中 1,204 个 | 真实终评 |
| A:成对题里行动力更高的一条得分更高 | 154 对中 149 对 | 合成 A 卷 |
| B:成对题里边界更清楚的一条得分更高 | 180 对中 179 对 | 两份合成 B 卷 |
| 与标签相差不超过 0.5(五维平均) | 87.3% | 工具包自检卷,195 道合成题;考过只说明接线对,不说明模型打得对 |
| 无法解析的输出 | 0 | 六份考卷 |
真实终评是三位已同意的内部员工的 453 轮对话,已假名化,所以这些不是一般人群上的准确率。合成考卷与训练同分布:考过只说明已知范围内没有退步,不说明能推广。
不是聊天机器人,不是诊断工具,也不是危机检测器。它不识别也不处理危机;它的分数,包括 W 在内,没有一个是危机信号。危机处理由先于模型运行的确定性代码负责:工具包的关键词闸门只是起点,不是完整筛查;面向消费者的心理健康部署,上游必须有独立机制,这是许可条件(LICENSE §3c)。
签字版预注册的 18 项检查,v10 过了 17 项。没过的那一项 G5a 当初是作为危机处理的替代指标设的;创始人看到结果后裁定,危机处理不在本模型里判定。按签字版预注册,判定是「拒收」;v10 由创始人决定发布。已报告的各项检查、考卷和方法,都在技术档案里。
合成对话由开放权重的教师模型打标,另有 440 轮假名化的内部员工对话(经本人同意)。外部用户的对话从不进入训练。训练消息里中文占 76.0%,英文 13.1%,中英混杂 10.9%;成绩没有按语言拆开测。
main 自 2026-10-06(UTC)起默认是 v10。固定版本请用 revision="f3869312f0222992c3eb2e938e78090de38eb80a"。gguf/hamo-score-0.6b-v10.q8.gguf;v9、v7、v6.1 的 GGUF 仍在 gguf/ 下。校验和与回退命令见技术档案。