Downloads · 30 days
0
litecreator/lite-llm
lite-llm is a machine learning model from litecreator. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Lite LLM is a deterministic, tiered-parameter, hierarchical sparse expert (HSER) language model runtime designed to scale from 1B → 1T parameters and beyond (up to quadrillion-scale parameter universes) while keeping…
Downloads · 30 days
0
Access
Public
Updated Apr 8, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.md5.6 KB · 79%
From the Hugging Face model README
Lite LLM is a deterministic, tiered-parameter, hierarchical sparse expert (HSER) language model runtime designed to scale from 1B → 1T parameters and beyond (up to quadrillion-scale parameter universes) while keeping active compute bounded per token.
The Github organization hosts the specification corpus, reference implementations, and operational tooling for building and deploying Lite LLM as an enterprise / reference-grade system.
Model optimization for the LiteCore Coherent Silicon Photonic Complex Multiply-Accumulate (CSP-cMAC) Unit Cell Hardware focuses on maximizing inference efficiency under tight memory and power constraints by combining compression, quantization, and memory aware execution. LiteCore is a fundamental photonic compute primitive purpose-built for large language model (LLM) inference at quadrillion-parameter scales. LiteCore leverages silicon-on-insulator (SOI) photonics to perform complex-valued multiply-accumulate operations at <1 fJ energy and 1–10 ps latency—representing 500–2,000× energy and 1,000–10,000× latency improvements over state-of-the-art electronic GPUs.
Lite LLM treats determinism as a first-class requirement:
Parameters are partitioned across storage tiers:
Only the TierSet for a request is eligible for routing; everything else has zero activation probability.
Routing is hierarchical:
Tier → Group → Expert
with bounded activation:
k_tier × k_group × k_expert experts per token per layer.
This enables extreme parameter scaling while keeping per-token compute predictable.
Lite LLM is not only a model architecture—it is a runtime system:
lite-llm-specs — Enterprise Runtime Engineering Specification Corpus (SPEC‑001…SPEC‑060)lite-llm-schemas — JSON/YAML schemas for manifests, telemetry, policieslite-llm-rfcs — Design proposals and evolution process (RFCs)lite-llm-runtime — Rust runtime (routing, caches, dispatch, TierSet engine)lite-llm-train — Training orchestration, checkpointing, determinism harnesslite-llm-kernels — Device kernels + safe wrappers (CUDA/HIP/Metal/CPU)lite-llm-comm — Transport abstraction (RDMA / NCCL / QUIC), collectiveslite-llm-storage — Shards, manifests, tier placement, streaming + prefetchlite-llm-cli — Operator CLI (inspect checkpoints, tier policies, telemetry)lite-llm-observability — Metrics exporters, dashboards, tracinglite-llm-deploy — Helm charts, Terraform modules, bare‑metal playbooksThe organization may not yet contain all repositories listed above; this is the intended long-term structure.
Start with:
The specs are written to be directly implementable:
Before performance optimization:
We welcome contributions in:
Please read:
CONTRIBUTING.md for workflow and standardsCODE_OF_CONDUCT.md for community expectationsSECURITY.md for vulnerability reportingLite LLM emphasizes:
See SECURITY.md to report vulnerabilities responsibly.
The specification corpus is the normative authority.
Changes to the corpus should go through the RFC process:
lite-llm-rfcsLite-LLM is distributed under the Dust Open Source License
license: other license_name: dosl-iie-1.0 license_link: https://github.com/lite-llm/lite-llm/raw/refs/heads/main/LICENSE
SECURITY.md