{"slug": "greenleaf-law-embed-tiny-a-compact-embedding-model-for-legal-domain-retrieval", "title": "GreenLeaf Law Embed Tiny: A Compact Embedding Model for Legal Domain Retrieval", "summary": "Researchers introduced GreenLeaf Law Embed Tiny, a 0.6B parameter embedding model for legal domain retrieval, achieving 75.11% on the Massive Legal Embedding Benchmark (MLEB) and 64.38% on MTEB(Law, v1). The model uses a two-stage training pipeline with knowledge distillation and domain-specific fine-tuning on 3.4 million query-passage pairs, including 150,000 human-curated samples, and supports BF16, INT8, and binary quantization for resource-constrained deployment.", "body_md": "arXiv:2608.24936v1 Announce Type: new\nAbstract: We present GreenLeaf Law Embed Tiny, a 0.6B parameter embedding model for legal domain retrieval. GreenLeaf-Tiny achieves 75.11% on the Massive Legal Embedding Benchmark (MLEB) and 64.38% on MTEB(Law, v1),demonstrating competitive performance among models under 1B parameters. Our approach combines a two-stage training pipeline that first distills knowledge from a larger teacher model into a compact student architecture, then applies domain-specific fine-tuning with hard negative mining; a carefully curated dataset of 3.4 million query-passage pairs, including 150,000 human-curated samples across diverse legal jurisdictions; and an efficient inference architecture supporting multiple quantization levels (BF16, INT8, binary) enabling deployment in resource-constrained environments. We provide detailed analysis of our training methodology, architectural choices, and comprehensive evaluation across legal retrieval tasks. Our results demonstrate that domain-specific training with high-quality data can improve performance for specialized domain applications", "url": "https://wpnews.pro/news/greenleaf-law-embed-tiny-a-compact-embedding-model-for-legal-domain-retrieval", "canonical_source": "https://arxiv.org/abs/2608.24936", "published_at": "2026-08-27 04:00:00+00:00", "updated_at": "2026-08-27 04:18:48.322512+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "natural-language-processing", "ai-research"], "entities": ["GreenLeaf Law Embed Tiny", "Massive Legal Embedding Benchmark", "MTEB(Law, v1)"], "alternates": {"html": "https://wpnews.pro/news/greenleaf-law-embed-tiny-a-compact-embedding-model-for-legal-domain-retrieval", "markdown": "https://wpnews.pro/news/greenleaf-law-embed-tiny-a-compact-embedding-model-for-legal-domain-retrieval.md", "text": "https://wpnews.pro/news/greenleaf-law-embed-tiny-a-compact-embedding-model-for-legal-domain-retrieval.txt", "jsonld": "https://wpnews.pro/news/greenleaf-law-embed-tiny-a-compact-embedding-model-for-legal-domain-retrieval.jsonld"}}