Downloads · 30 days
8
28% of all-time downloads
mki0809/geosearch-reranker
geosearch-reranker is a machine learning model from mki0809. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-4.0.
CatboOst YetiRank listwise reranker, the last stage of a multilingual toponym search pipeline:
Downloads · 30 days
8
28% of all-time downloads
All-time downloads
29
Public
Parameters
118M
490 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors471 MB · 96%
From the Hugging Face model README
CatboOst YetiRank listwise reranker, the last stage of a multilingual toponym
search pipeline:
text
-> GLiNER NER
-> char-n-gram BM25 retrieval
-> CatBoost reranker (this model)
-> ranked GeoNames places
It reorders a fixed candidate set; it does not choose one. Retrieval ends by ranking on BM25 score with population as the tiebreak and cutting to top-k, and the model only permutes those survivors.
| system | RR | P@1 | R@5 | R@10 | R@25 | R@50 |
|---|---|---|---|---|---|---|
| baseline | 0.8692 | 0.8333 | 0.8357 | 0.8813 | 0.9429 | 0.9578 |
| rerank | 0.9147 | 0.8833 | 0.8940 | 0.9239 | 0.9543 | 0.9578 |
| ideal | 0.9833 | 0.9833 | 0.9341 | 0.9461 | 0.9566 | 0.9578 |
The loader validates feature_names_ against this exact list and refuses a
model that does not match, so a checkpoint trained on a different feature set
degrades to retriever order instead of scoring garbage. Published so that
failure is diagnosable by someone who did not train it:
city_entitiescountry_entitiesadmin1_entitiesdocumentlog_populationretriever_scoreretriever_rankcity_exact_matchcity_substr_matchcity_ngram_containmentcity_token_covercountry_matchhas_country_spanadmin1_substr_matchadmin1_ngram_containmenthas_admin1_spanText features (CatBoost text_features) are the NER spans split by type —
city_entities / country_entities / admin1_entities — plus the candidate
document.
The typed split does not by itself match a span against a candidate.
CatBoost builds a bag of words per text column independently, so it never sees
the intersection of city_entities and document. A model trained on the
split alone gave all three entity buckets 1.5% of its importance combined and
scored below the retriever's own order. The overlap is therefore computed
explicitly — the *_match / *_containment / *_cover features above — with
has_country_span / has_admin1_span so a real mismatch (0.0) stays
distinguishable from "no comparison was possible" (-1.0).
The document a candidate is scored against is three lines:
<name> | <name> | <name>
<country in English>
<admin1 region in English>
The ' | ' separator is load-bearing. Joined by a space instead,
a place's spellings collapse into one pseudo-name whose n-grams are only ~1/n
covered by a span naming it once — which turns the containment feature into an
inverse-popularity signal. Separated, each spelling is matched alone and the
best wins.
use_gold_entities: FalseThat flag has to be false for a servable model. true trains on the query
dataset's gold spans instead of mined NER output — a useful ablation ("how good
would this be if NER were perfect?") that breaks train/serve parity by design,
because online the reranker always receives GLiNER spans. A model trained with
it set must not be served, and this line is on the card so that cannot happen
silently.
Split by query, not by place: a query is the ranking group, so splitting on geonameid would tear one query's candidates across train and test — leaking the query and leaving positive-less test groups.
admin1_name exists in English only (admin1CodesASCII.txt), so the admin1
features work for en/tr and read 0 for ru/zh. Regions are
feature_class='A' and the ETL loads only 'P', so no localised region names
were ever ingested.
Derived from GeoNames, licensed CC BY 4.0.
Modifications made to the source data:
feature_class = 'P')