Downloads · 30 days
30
34% of all-time downloads
OptGear/Opt.Gear-1M
Opt.Gear-1M is a text generation model from OptGear. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as cc-by-nc-sa-4.0.
<img width="1000px" src="./OptAIxHuggingFace03.png"
Downloads · 30 days
30
34% of all-time downloads
All-time downloads
87
Public
Parameters
1M
2.3 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.1 MB · 84%
From the Hugging Face model README
[!Note] This repository contains model weights for Opt.Gear-1M, a Tiny Language Model (TLM) designed to run on Micro-Controller Units (MCUs), together with artifacts for embedded deployment. <!-- TODO: 리포 실제 구성물(가중치 포맷, 임베디드 빌드 아티팩트) 확정 후 수정 -->
Opt.Gear-1M is not a general-purpose chatbot. It is trained for a narrow embedded-control setting — mapping short natural-language commands to structured board-control intents — and is intended for deeply embedded systems such as sensors, controllers, small robots, and offline human-machine interfaces.
For general-purpose on-device text generation, see Opt.Gear-270M and Opt.Gear-1B. <!-- TODO: 링크 확인 -->
The goal of Opt.Gear-1M is not to shrink Opt.Gear-270M/1B further, but to actually run an auto-regressive language model on a microcontroller — an environment far more constrained than mobile NPUs or server GPUs. Deployed directly on the ARM Cortex-M7 core of an STM32H747I-DISCO board with 4-bit weights and FP32 activations (W4A32), Opt.Gear-1M generates at:
20 tokens/s on a 400MHz STM32 Cortex-M7 — about 50ms per token.
Measured on the actual MCU with an embedded C runtime, not a host-side simulator. To our knowledge, Opt.Gear-1M is the first generative language model to achieve 20 TPS with W4A32 quantization on the ARM Cortex-M7 CPU of the STM32H747I-DISCO.
For more details, please refer to our tech report and blog post. <!-- TODO: 링크 연결 -->
stm32_cmd_synth (no separate web-scale pre-training phase)Opt.Gear-1M follows a different training path from Opt.Gear-270M/1B. It is trained on stm32_cmd_synth, a synthetic command dataset of 50,000 prompt-response pairs: each prompt is a short natural-language command for controlling the STM32 board, and each target response is a compact JSON-style control intent. The action space covers LCD, LED, and camera control. This design intentionally favors predictable command generation over broad open-domain language ability, matching the constraints of MCU deployment.
# Illustrative example — TODO: 실제 데이터 포맷으로 교체
User: turn on the red LED and show "hello" on the screen
Model: {"led": {"color": "red", "state": "on"}, "lcd": {"text": "hello"}}
All numbers are measured on the actual device. The firmware is instrumented with HAL_GetTick; throughput is measured from the Cortex-M7 execution path and reflects the model-compute portion of auto-regressive generation. Loadable section sizes are measured from the final STM32 ELF with arm-none-eabi-size.
| Item | Measured or configured value |
|---|---|
| Board | STM32H747I-DISCO |
| MCU / core | STM32H747XIH6 / ARM Cortex-M7 |
| Configured CPU clock | 400MHz |
| Runtime | Embedded C (no ONNX / X-CUBE-AI dependency) |
| Quantized format | W4A32 (q4 weights, q16 embedding/LM head) |
| Deployment sequence length | 160 tokens |
| Maximum new tokens | 40 tokens |
| Measured throughput | 20 tokens/s |
| Per-token latency | 50ms/token |
.text / .rodata | 1,596,400 bytes |
.data | 532 bytes |
.bss | 444,404 bytes |
| Total static RAM (incl. reserved heap/stack) | 449,032 bytes |
A few notes on reading these numbers:
.text/.rodata (≈1.5MB) is not the pure weight size: it includes the embedded runtime code, quantized weights, 16-bit embedding/LM head, lookup tables, scales, and constants. For MCU models, the final ELF — not the theoretical parameter count — is the meaningful footprint measure.The measured binary uses an embedded C runtime generated for STM32CubeIDE and runs without depending on an ONNX graph or the X-CUBE-AI generated network. The deployment build uses a 160-token sequence length and a 40-token generation limit, keeping the auto-regressive cache and workspace bounded on the MCU while preserving enough context for short instructions and embedded-control prompts.
<!-- TODO: 임베디드 빌드/플래싱 가이드, 예제 펌웨어 링크 추가 -->Potential applications at this speed and footprint:
Stay in the trained domain: The model is post-trained for short English board-control commands with JSON-style outputs. Out-of-domain prompts (open-ended questions, long-form generation, non-English input) will not produce reliable results.
Respect the deployment limits: The model configuration targets a 512-token maximum context, and the reference deployment build uses 160-token sequences with a 40-token generation cap. Longer sequences increase the SRAM-resident cache and workspace.
Budget by ELF, not parameter count: When adapting the runtime or retraining the model, verify Flash/SRAM budgets with arm-none-eabi-size on the final ELF — runtime code, lookup tables, and scales share the Flash with the weights.
Predictability over coverage: On embedded systems, predictable memory use and stable interactive latency matter more than long-context benchmark performance. The fixed-size ConvKV state exists precisely to keep decoding state independent of context length.
Opt.Gear-1M is not a general-purpose chatbot. The ~1M parameter scale and 2,048-token vocabulary impose clear limits: it targets stable generative capability under extremely small flash, SRAM, and compute budgets, not broad benchmark coverage. The model prioritizes compact English generation, short-form command following, and predictable structured outputs. For general text generation on mobile and edge devices, use Opt.Gear-270M or Opt.Gear-1B.
If you find our work helpful, feel free to give us a cite.
@misc{optgear2026,
title = {{Opt-Gear} Technical Report},
author = {{Opt.Gear Team}},
year = {2026},
url = {https://huggingface.co/OptGear}
}
Correspondence: contact@opt-ai.kr · Hugging Face: huggingface.co/OptAI