Evaluating RAG for French immigration law: a benchmark and baseline study A new benchmark and baseline study evaluating retrieval-augmented generation (RAG) for French immigration law finds that retrieval improves administrative guidance at both model scales tested. Comparing a parametric LLM baseline against dense retrieval augmentation with Qwen3.5-9B and Qwen3.5-27B on 52 annotated synthetic profiles, the study shows retrieval most notably improves permit-type accuracy. The results confirm that retrieval grounding is important for more reliable administrative guidance in this domain. arXiv:2607.24449v1 Announce Type: cross Abstract: International recruitment in France requires navigating a layered legal framework absent from existing legal AI benchmarks. We present a publicly available benchmark and first comparative evaluation for this domain, covering permit-type recommendation, required-document retrieval, and legal citation coverage. Comparing a parametric LLM baseline against dense retrieval augmentation at two model scales Qwen3.5-9B and -27B on 52 annotated synthetic profiles, we find that retrieval improves administrative guidance at both scales, most notably permit-type accuracy. Our results confirm that retrieval grounding is important for more reliable administrative guidance in this domain, and motivate further investigation of hybrid retrieval strategies.