Downloads · 30 days
0
irinaqqq/lexir
lexir is a sentence similarity model from irinaqqq. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as other.
LexIR is a bilingual Russian/Kazakh legal AI system that combines fine-tuned semantic retrieval, FAISS vector search, deterministic rule evaluation, and Azure OpenAI Assistant orchestration.
Downloads · 30 days
0
Access
Public
Updated Sep 1, 2026
Repo size
4.5 GB
Likes
0
Public
Click a slice to open those files.
.safetensors2.2 GB · 70%
From the Hugging Face model README
LexIR is a bilingual Russian/Kazakh legal AI system that combines fine-tuned semantic retrieval, FAISS vector search, deterministic rule evaluation, and Azure OpenAI Assistant orchestration.
The project is designed around a simple principle: the language model should not answer legal questions from its general knowledge. Instead, it retrieves relevant legal provisions from an explicitly connected corpus and generates answers grounded in those sources.
The architecture is corpus-agnostic and can be adapted to different legal collections. The current demo configuration uses Kazakhstan legal data.
The retrieval pipeline was evaluated on a held-out benchmark with no positive-example overlap between the training and evaluation sets.
| Metric | Result |
|---|---|
| Aligned RU/KZ legal clauses | 17.5K |
| Legal sources | 6 |
| Evaluation queries | 699 |
| Fine-tuned model Recall@10 | 0.355 |
| Improvement over base MPNet / LaBSE | 3.8–4.2× |
The final retriever uses a multilingual MPNet-based Sentence Transformer fine-tuned with Multiple Negatives Ranking Loss.
The project also includes baseline evaluation against non-fine-tuned MPNet and LaBSE models.
flowchart LR
U[User] --> API[FastAPI]
API --> AOAI[Azure OpenAI Assistant]
AOAI -->|Tool call| RET[Retrieval Tool]
RET --> RU[RU FAISS Index]
RET --> KZ[KZ FAISS Index]
RU --> MERGE[Merge and Rank]
KZ --> MERGE
MERGE --> SOLVER[Deterministic Rule Solver]
SOLVER --> AOAI
AOAI --> ANSWER[Grounded Answer + Citations]
ANSWER --> U
API --> TRACE[SQLite / JSONL Trace Store]
RET --> TRACE
SOLVER --> TRACE
The runtime combines LLM orchestration with local deterministic components.
Azure OpenAI manages the dialogue and tool-call lifecycle, while retrieval and rule evaluation remain under application control.
A parser/ETL pipeline extracts legal provisions and converts them into a normalized bilingual representation.
The current project includes a Kazakhstan legal corpus with aligned Russian and Kazakh text.
The pipeline prepares query-to-relevant-clause training and evaluation examples.
Data validation includes consistency checks and protection against train/evaluation leakage.
src/train_biencoder.py fine-tunes a multilingual MPNet-based Sentence Transformer on query-to-clause pairs.
The training pipeline uses:
src/build_index.py generates independent RU and KZ FAISS indexes for different model variants.
The project supports evaluation of:
At runtime, a user query is searched against both Russian and Kazakh indexes.
The retrieval layer:
Azure OpenAI Assistants manage conversation threads and invoke local application tools.
The assistant uses retrieved legal provisions rather than relying on unsupported model knowledge.
Tool execution, retrieval results, and reasoning metadata are persisted for later inspection.
For cases involving explicit numeric or boolean legal constraints, LexIR can evaluate structured conditions outside the LLM.
This separates deterministic decision logic from generative reasoning and makes the resulting workflow easier to inspect and audit.
artifacts/
models/
finetuned_mpnet/
indexes/
<alias>/
ru.faiss
kz.faiss
ru_meta.jsonl
kz_meta.jsonl
reports/
data/
clauses_constitution_ru_kz.jsonl
legal_assistant_train.jsonl
legal_assistant_test.jsonl
data_parser/
adilet_zan_parser.py
site/
backend/
app.py
assistant/
assistant_create.py
assistant_edit.py
assistant_info.py
demo_assistant.py
frontend/
index.html
app.js
styles.css
src/
build_index.py
train_biencoder.py
evaluate.py
plot_eval.py
validate.py
demo_cli.py
api.py
Legal clauses are stored as JSONL records containing bilingual text and structured metadata.
{
"id": "KZ.CONST.1995:ART18:PAR2:cl1",
"text": "Russian legal provision...",
"text_kz": "Қазақ тіліндегі құқықтық норма...",
"meta": {
"doc_id": "KZ.CONST.1995",
"article_number": "18",
"paragraph_number": 2,
"article_title_ru": "...",
"article_title_kz": "...",
"source_ru": "...",
"source_kz": "..."
}
}
The data format is intentionally separated from the retrieval implementation so that another legal corpus can be substituted without redesigning the entire system.
The default retrieval model is the fine-tuned Sentence Transformer.
Clause embeddings are normalized and indexed using FAISS IndexFlatIP, which provides cosine-similarity search over normalized vectors.
The retrieval pipeline accepts configurable parameters such as:
top_kEach result contains:
Russian and Kazakh retrieval can be executed independently and combined into a single ranked result set.
This allows LexIR to retrieve semantically relevant provisions even when the query language and the most useful corpus representation differ.
The runtime implementation supports concurrent RU/KZ retrieval and falls back to sequential execution if parallel retrieval fails.
The retriever is trained using query-to-positive-clause examples.
The workflow includes:
The final multilingual MPNet retriever achieved Recall@10 = 0.355, representing a 3.8–4.2× improvement over the evaluated base MPNet and LaBSE configurations on the 699-query held-out benchmark.
The web application uses Azure OpenAI Assistants for conversation orchestration.
The assistant can invoke the local LexIR retrieval tool and construct an answer from the returned legal provisions.
Assistant configuration is located under:
site/backend/assistant/
The Assistant ID can be supplied through configuration or generated using the provided setup script.
LexIR stores structured traces of runtime interactions.
Depending on configuration, traces include:
SQLite is used for persistent runtime history, with JSONL available as an additional audit representation.
This makes it possible to inspect how an answer was produced instead of treating the LLM response as an opaque result.
python -m venv .venv
source .venv/bin/activate
pip install -r site/backend/requirements.txt
Create the required environment variables:
AZURE_OPENAI_API_KEY=...
AZURE_OPENAI_VERSION=...
AZURE_OPENAI_ENDPOINT=...
Optional configuration:
ASSISTANT_ID=...
AZURE_OPENAI_ASSISTANT_MODEL=...
The backend expects the fine-tuned model and FAISS indexes under:
artifacts/models/finetuned_mpnet/
artifacts/indexes/finetuned/ru.faiss
artifacts/indexes/finetuned/kz.faiss
artifacts/indexes/finetuned/ru_meta.jsonl
artifacts/indexes/finetuned/kz_meta.jsonl
If the artifacts have not been generated:
python data_parser/adilet_zan_parser.py
python src/train_biencoder.py
python src/build_index.py
python site/backend/assistant/assistant_create.py
uvicorn app:app --app-dir site/backend --host 0.0.0.0 --port 8000
Open:
http://localhost:8000
docker compose up --build
The application is exposed on port 8000.
Run local semantic retrieval without Azure OpenAI:
python src/demo_cli.py
Run the Assistant-based CLI:
python site/backend/assistant/demo_assistant.py
Run dataset validation:
python src/validate.py
Run retrieval evaluation:
python src/evaluate.py
Generate evaluation plots:
python src/plot_eval.py
Reports and figures are stored under:
artifacts/reports/
artifacts/reports/figures/
The architecture is not tied to a single legal document.
To use another corpus:
Backend
AI / ML
Data & Evaluation
Infrastructure
LexIR is a research and engineering prototype.
Current limitations include:
LexIR is intended for research, experimentation, and software engineering demonstrations.
It does not provide legal advice.
Legal conclusions should be verified against official and current legal sources.