Marathi-XRetrieve: Cross-Lingual Retrieval Benchmark
Independent ResearchObjective: Quantify and address performance gaps in cross-lingual retrieval for Marathi, a low-resource Indic language with over 80 million speakers.
Methodology: Developed a reproducible benchmark with 1,000 Marathi queries over English/Hindi Wikipedia. Evaluated BGE-M3 zero-shot performance and applied contrastive fine-tuning on 1,000 query-chunk pairs.
Key Results:
- Baseline cross-lingual gap: 7.93 percentage points in Top-1 accuracy (Marathi vs. English queries)
- Fine-tuning reduced retrieval error by 13% relative for Marathi queries (77.36% to 80.20% Top-1 accuracy)
- Gap reduction: 47.4% with no significant degradation on English performance (-0.92 pp)
- Comprehensive evaluation using MRR, NDCG@k, and MAP metrics with bootstrap confidence intervals