Running Mistral offline: the tokenizer is on disk, but the library can't find it A developer identified and fixed a bug in mistral-common, Mistral AI's open-source request-preparation library, that caused offline tokenizer loading to fail with a FileNotFoundError when models were stored outside the default Hugging Face cache. The fix, merged by a Mistral AI maintainer in PR #349, makes the file-listing step honor the user-supplied cache_dir the same way hf_hub_download does, so air-gapped deployments with a dedicated model disk can load tokenizers without network access. The developer notes Transformers and vLLM likely inherit the same issue offline, though that was not tested end to end. An inference server cut off from the Internet. A dedicated disk for models. Mistral's tokenizer downloaded in advance, sitting in its folder. And at startup: FileNotFoundError: No local files found for the repo ID mistralai/Mistral-7B-v0.1 and revision None. The file is there. The library looks somewhere else. I found this bug in mistral-common https://github.com/mistralai/mistral-common , Mistral AI's open-source library that prepares requests for its models. The fix was reviewed, approved and merged by a Mistral AI maintainer PR 349 https://github.com/mistralai/mistral-common/pull/349 , issue 348 https://github.com/mistralai/mistral-common/issues/348 . Teams running Mistral models without Internet access that store their models outside the default Hugging Face cache . That is exactly the enterprise setup: air-gapped network for security, corporate proxy, a dedicated model disk on the GPU server. Two situations trigger it: local files only=True ; HF HUB OFFLINE=1 and the library falls back to local files on its own. With the default cache, or with network access, nothing happens. That is why the bug stays invisible in development and shows up on deployment day. Offline, loading the tokenizer chains two functions: list local hf repo files lists the cached files to find the tokenizer tekken.json . It HF HUB CACHE , and had no cache dir parameter. hf hub download then fetches the file. It does receive before: the listing ignores the folder chosen by the user repo cache = Path huggingface hub.constants.HF HUB CACHE / ... If the tokenizer only lives in cache dir , the listing comes back empty and the function gives up before even trying to read the file. The reverse case fails too, one step later. Both steps now read the same folder: cache dir if given, otherwise the default cache, exactly the rule hf hub download follows. cache root = Path cache dir if cache dir is not None else Path huggingface hub.constants.HF HUB CACHE repo cache = cache root / huggingface hub.constants.REPO ID SEPARATOR.join "models", repo id.split "/" Path , full offline load, fallback after a network error, The maintainer asked for a single change and applied it himself: removing a code comment. The code itself did not change. If you are affected and cannot upgrade mistral-common yet: point HF HUB CACHE at the same folder as your cache dir . Both steps then read the same place. From reading their source code, Transformers cache dir and vLLM --download-dir with tokenizer mode=mistral pass that folder down to mistral-common, so they are probably affected offline too. I did not test that end to end. Deploying Mistral models on an air-gapped network? Which other traps have you hit? Let's talk, or get your own agents audited: LinkedIn https://www.linkedin.com/in/zakariakhchiche/ · Malt https://www.malt.fr/profile/zakariakhchiche · Medium https://medium.com/@ZKHCHICHE · Website https://zakariakhchiche.github.io/ Want your team to build agents like these? I run a hands-on Copilot Studio training in French, with Spar-x, Qualiopi-certified, eligible for OPCO funding in France : https://zakariakhchiche.github.io/formation-copilot-studio/ https://zakariakhchiche.github.io/formation-copilot-studio/ and a free AI Act article 4 kit: https://zakariakhchiche.github.io/kit-ai-act/ https://zakariakhchiche.github.io/kit-ai-act/