Downloads Β· 30 days
0
shilpesh619/Leonet
Leonet is a machine learning model from shilpesh619. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This repository contains a demonstration of LeoNet, a biologically inspired transformer architecture that unifies language understanding with motor action execution.
Downloads Β· 30 days
0
Access
Public
Updated Jul 7, 2025
Repo size
196 MB
Likes
0
Public
Click a slice to open those files.
.pth196 MB Β· 99%
From the Hugging Face model README
This repository contains a demonstration of LeoNet, a biologically inspired transformer architecture that unifies language understanding with motor action execution.
LeoNet transforms typed natural language commands (e.g., "move left", "scroll down") into motor control vectors capable of moving the mouse pointer across the screen.
| Filename | Description |
|---|---|
Leonet_model.py | The LeoNet transformer architecture definition. |
leonet_command_vision_500.jsonl | A small sample dataset of command-to-motor pairs (8x128 motion vectors). |
leonet_train_pipeline.py | The model training script (text + motor prediction). |
Leonet_inference.py | Inference script for typing a command and observing real mouse movement. |
LeoNet introduces a novel architecture inspired by the brainβs ability to convert cognition into action. The model accepts language commands and outputs two parallel predictions:
The motor stream is trained to directly map commands like "move right" into cursor deltas like [+200, 0], simulating brain-to-muscle control.
This prototype bridges the gap between natural language processing and motor control, setting the foundation for fully embodied LLM-based agents.
Make sure you have Python 3.9+ and install dependencies:
pip install torch pyautogui pyttsx3
Ensure the following files are in the same folder:
Leonet_model.pyleonet_command_vision_500.jsonlleonet_train_pipeline.pyLeonet_inference.pyYou can train the LeoNet model on the dataset using:
python leonet_train_pipeline.py
After training, this will create:
leonet_demo.pth
python Leonet_inference.py
Youβll see:
β¨οΈ Type a command to move (e.g., 'move left', 'scroll down')
π Type 'quit' to exit.
Command:
Type:
move left
LeoNet will:
dx = -200, dy = 0LeoNetβs transformer has:
Motor output is trained using MSE loss to match known action deltas from data like:
{
"input_ids": [13, 14, 17, 4, 26, 26, 26, 26],
"target_ids": [14, 17, 4, 26, 26, 26, 26, 0],
"motor_output": [[-200.0, 0.0], [0.0, 0.0], ..., [0.0, 0.0]]
}
| Model | Parameters (M) | FLOPs (GFLOPs) | Inference Latency (ms) | Motor Output Support | Dual Output (Lang + Motor) | Use Case Suitability |
|---|---|---|---|---|---|---|
| LeoNet | 29.55 | 0.818 | 54.5 | β 1024-dim | β Yes | Real-world dual-task |
| TinyBERT | 14.5 | 1.3 | ~66 | β None | β No | Lightweight NLP |
| DistilBERT | 66 | 3.8 | ~115 | β None | β No | Faster BERT |
| BERT Base | 110 | 12.0 | ~150 | β None | β No | Deep language tasks |
| GPT-2 Small | 124 | 15.5 | ~180 | β None | β No | Generative tasks |
LeoNet is designed not just for research, but for real-world deployment in low-resource environments like Raspberry Pi and NVIDIA Jetson platforms.
| Feature | Supported |
|---|---|
| Dual Output (Language + Motor) | β Yes |
| Low FLOPs (~0.818 GFLOPs) | β Yes |
| Compact Model (~29.5M Parameters) | β Yes |
| Inference Latency (~54 ms) | β Yes |
| Suitable for Real-Time Robotics | β Yes |
| Runs on Jetson Nano / Xavier / Pi 4 | β Yes |
Input:
"move cursor right"
LeoNet Output:[0.12, 0.00, ..., -0.04]
Action: Simulated or real motor movement (e.g., servo, cursor, wheel)
LeoNet bridges language understanding and physical action, making it ideal for embodied agents, GUI automation, and embedded AI.
Silpeshkumar Jitendrabhai Patel.
LeoNet: A Brain-Inspired Transformer for Dual Cognitive and Motor Output in Real-World Environments.
TechRxiv Preprint, 2025.
This project is released under the CC BY 4.0 License. You are free to use, modify, and distribute with attribution.
This project demonstrates a proof of concept for building LLM-based motion agents that "think and act" in real-world environments. It combines: