Downloads · 30 days
9
16% of all-time downloads
venky1/entrofrs-small
entrofrs-small is a machine learning model from venky1. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
[](LICENSE) [](https://www.python.org/downloads/)
Downloads · 30 days
9
16% of all-time downloads
All-time downloads
55
Public
Repo size
132 KB
Likes
0
Public
Click a slice to open those files.
.npz132 KB · 62%
From the Hugging Face model README
EntroFRS is a lightweight, offline, question-aware Employee/FRS data summarization system.
It runs entirely locally after the first setup — no cloud APIs, no paid services, no internet required at inference time.
EMPLOYEE/FRS DATA
↓
INPUT FORMAT NORMALIZATION
(JSON / Key-Value / Pipe / Natural Language / Mixed)
↓
QUESTION/INTENT UNDERSTANDING
(Rules + INT8 Hashed-Feature Linear Classifier)
↓
FACT EXTRACTION & GROUNDING
(Schema-aware deterministic parsing)
↓
GROUNDED RENDERER
(Controlled English generation from extracted facts)
↓
QUESTION-SPECIFIC PROFESSIONAL SUMMARY
| Criterion | Generative LLM | This System |
|---|---|---|
| Model size < 25 MB | Very difficult | ✅ ~150 KB model assets |
| Factual accuracy | Hallucination risk | ✅ Copies source facts |
| Question relevance | Requires instruction tuning | ✅ Intent-based selection |
| CPU inference speed | Slow | ✅ Sub-millisecond |
| Offline capability | Needs large download | ✅ Tiny, fully offline |
| No paid APIs | Often needs cloud | ✅ 100% local |
# Install from source
pip install -e .
# Or with API support
pip install -e ".[api]"
# 1. Generate synthetic training data
python -m entrofrs_llm.tools dataset
# 2. Train the intent classifier
python -m entrofrs_llm.tools train
# 3. Export INT8 quantized model
python -m entrofrs_llm.tools export
# 4. Evaluate on 200 test cases
python -m entrofrs_llm.tools evaluate
# 5. Measure model size
python -m entrofrs_llm.tools measure
# 6. CPU benchmark
python -m entrofrs_llm.tools benchmark
from entrofrs_llm import EntroFRSModel
model = EntroFRSModel.load()
# JSON input
response = model.summarize(
{
"employee_name": "Robert Anderson",
"role": "DevOps Engineer",
"manager": "Karen White",
"work_date": "August 24, 2026",
"attendance": "Late",
"working_hours": "8.50",
"overtime": "0.50"
},
question="What is Robert's role?"
)
print(response)
# → "The employee's role is DevOps Engineer."
# Key-value text input
response = model.summarize(
"Employee: Robert Anderson\nRole: DevOps Engineer\nManager: Karen White",
question="Who is the manager?"
)
print(response)
# → "Manager: Karen White."
# Missing information
response = model.summarize(
{"employee_name": "Robert Anderson", "role": "DevOps Engineer"},
question="What was the overtime?"
)
print(response)
# → "The requested information is not available in the provided employee data."
export ENTROFRS_API_KEY="your-secret-key"
uvicorn entrofrs_llm.api:app --host 127.0.0.1 --port 8000
curl http://127.0.0.1:8000/summarize \
-H 'Content-Type: application/json' \
-H 'X-API-Key: your-secret-key' \
-d '{
"data": {"employee_name": "Robert Anderson", "role": "DevOps Engineer"},
"question": "What is the employee role?"
}'
Endpoints:
GET /health — Health checkGET /info — Model informationPOST /summarize — Question-aware summarization| Format | Example |
|---|---|
| JSON | {"employee_name": "Robert", "role": "Engineer"} |
| Key-Value | Employee: Robert\nRole: Engineer |
| Pipe (header) | Employee | Role\nRobert | Engineer |
| Natural Language | Robert is an Engineer reporting to Karen. |
| Mixed | Employee: Robert\nThe role is Engineer. |
| Field | Aliases |
|---|---|
| employee_name | employee, name, emp name |
| employee_id | emp id, staff id |
| role | designation, job title |
| department | dept |
| manager | reporting manager, reports to |
| location | office location |
| joining_date | date joined, join date |
| team | team name |
| work_date | date |
| attendance | attendance status |
| present_days | days present |
| absent_days | days absent |
| late_days | days late |
| attendance_percentage | attendance percent, attendance pct |
| leave | leave information, leave info |
| clock_in | clock in, check in |
| clock_out | clock out, check out |
| working_hours | work hours, hours worked |
| effective_hours | effective work hours |
| overtime | overtime hours, ot |
| tasks | task |
| completed_tasks | completed task, tasks completed |
| project | projects, project name |
| task_status | status |
| approval_status | approval |
| comments | comment, notes |
| technologies | technology stack, tech stack |
| worklogs | worklog, work logs, time logs |
After the first load, the model is cached at:
~/.cache/entrofrs/<content-hash>/ENTROFRS_CACHE environment variableSubsequent loads use the cached version. No internet required after initial setup.
The evaluation system measures:
entrofrs/
├── pyproject.toml # Build configuration
├── requirements.txt # Dependencies
├── LICENSE # MIT License
├── README.md # This file
├── entrofrs_llm/
│ ├── __init__.py # Public API
│ ├── schema.py # Field definitions, tokenizer, features
│ ├── core.py # Normalization, intent routing, rendering
│ ├── artifacts.py # Model caching and verification
│ ├── api.py # FastAPI endpoints
│ ├── data.py # Dataset generation
│ ├── tools.py # CLI: train, export, evaluate, benchmark
│ └── model/ # Exported INT8 model assets
│ ├── intent.int8.npz
│ ├── config.json
│ ├── manifest.json
│ └── LICENSE
├── datasets/ # Generated training/test data
├── artifacts/ # FP32 training checkpoints
└── reports/ # Evaluation and benchmark results
MIT — See LICENSE