Downloads · 30 days
227
100% of all-time downloads
Ayushi054/CourtInsight-BNS-Encoder
CourtInsight-BNS-Encoder is a sentence similarity model from Ayushi054. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as apache-2.0.
CourtInsight BNS Encoder is a BNS-specific semantic embedding model developed for the CourtInsight legal-information retrieval system.
Downloads · 30 days
227
100% of all-time downloads
All-time downloads
227
Public
Parameters
137M
3.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors547 MB · 100%
From the Hugging Face model README
CourtInsight BNS Encoder is a BNS-specific semantic embedding model developed for the CourtInsight legal-information retrieval system.
The model is designed to retrieve relevant provisions of the Bharatiya Nyaya Sanhita, 2023 (BNS) from natural-language descriptions of legal situations.
The retrieval approach is semantic and does not use keyword matching, BM25 retrieval, manually defined keyword lists, or hard-coded section mappings.
The model is not a legal-advice system and should not be used to determine legal liability, make judicial decisions, or replace qualified legal professionals.
nomic-ai/nomic-embed-text-v1.5
The model is compatible with Sentence Transformers.
For queries, use the prefix:
search_query:
For documents, use:
search_document:
The training corpus was constructed from a structured dataset covering 358 BNS 2023 sections.
The dataset contains statutory information including section number, section title, operative text, explanation, illustration, legal elements, and element summaries.
Semantic hard-negative mining was performed using the base Nomic embedding model.
The V3 training dataset contained 10,545 query-positive-negative triplets.
Training used TripletLoss with cosine distance.
| Parameter | Value |
|---|---|
| Base model | nomic-ai/nomic-embed-text-v1.5 |
| Loss | TripletLoss |
| Distance | Cosine distance |
| Triplet margin | 0.5 |
| Epochs | 5 |
| Effective batch size | 16 |
| Learning rate | 2e-5 |
| Maximum sequence length | 512 |
| Precision | FP16 |
15 manually constructed legal scenarios:
These results use the complete CourtInsight retrieval pipeline.
This benchmark is statute-derived and should not be interpreted as real-world user-query accuracy.
This benchmark is synthetic/title-derived and should not be interpreted as real-world legal accuracy.
User Query
→ CourtInsight BNS Encoder
→ Top-10 Semantic Candidates
→ Cross-Encoder Reranking
→ Nomic + Cross-Encoder Score Fusion
→ Top-3 BNS Provisions
The retrieval pipeline does not use keyword matching, BM25, or hard-coded section mappings.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer(
"Ayushi054/CourtInsight-BNS-Encoder",
trust_remote_code=True
)
query = "A person knowingly uses a forged document as genuine."
embedding = model.encode(
"search_query: " + query,
normalize_embeddings=True
)
Apache-2.0
CourtInsight is an experimental/research legal-information retrieval system. It is intended to help users locate potentially relevant BNS provisions. It does not provide legal advice and should not be treated as a substitute for a qualified legal professional or authoritative legal source.