503 Service Unavailable hitting multiple major Image-to-3D Spaces (TripoSR, InstantMesh, LGM) via Gradio Client Four major Hugging Face Image-to-3D Spaces — TripoSR, InstantMesh, CRM, and LGM — are showing Runtime errors from four distinct latent startup failures rather than one shared cause, according to a developer analysis of the public Spaces. The proposed minimal fixes are adding onnxruntime to TripoSR and InstantMesh (with python_version 3.10.13 pinned in InstantMesh's README), replacing CRM's dead stabilityai/stable-diffusion-2-1-base scheduler source with sd2-community/stable-diffusion-2-1-base, and adding nvidia-cuda-runtime-cu11 to LGM plus preloading libcudart.so.11.0 before importing its compiled extension. The analysis states the failures are deterministic in each repo's own code path, so the patches are required even if a platform-side restart, unpause, or rebuild triggered the cold starts that exposed them. Based on the issues I identified while modifying my own Spaces, the following are the minimum necessary concrete fixes: All four currently show Runtime error , but they do not fail for one shared code reason. What happened is closer to this: a platform-side event likely forced cold starts, and those cold starts exposed four different latent startup problems. So the right repair strategy is not “apply one global workaround,” but “make each repo boot cleanly in the current Spaces environment with the smallest justified diff.” The four buckets are: missing onnxruntime after rembg import for TripoSR and InstantMesh, a dead upstream repo id for CRM, and a native CUDA runtime mismatch for LGM. HF’s current Spaces config still supports explicit python version pinning, and current ZeroGPU docs list 3.10.13 and 3.12.12 as supported Python versions, so version pinning is still part of the stabilization story. Hugging Face https://huggingface.co/spaces/stabilityai/TripoSR/blob/main/README.md The key principle is this: even if the trigger was a restart, unpause, rebuild, scheduler issue, or temporary API problem, these fixes are still needed because each current repo has a deterministic startup failure in its own code path. In other words, even if the platform caused the failure to become visible, the repo still has to be made bootable. That is why I would keep the patches minimal and specific instead of doing large framework upgrades first. pip has become stricter over time, but none of the four currently exposed failures are primarily “requirements syntax” bugs. They are startup dependency and binary/runtime mismatches. pip https://pip.pypa.io/en/latest/news/ 1. TripoSR : add onnxruntime . 2. InstantMesh : add onnxruntime , and pin python version: 3.10.13 in README. 3. CRM : replace the dead stabilityai/stable-diffusion-2-1-base scheduler source with sd2-community/stable-diffusion-2-1-base . 4. LGM : add nvidia-cuda-runtime-cu11 , then preload libcudart.so.11.0 before importing the compiled extension. This is the smallest targeted fix for the current public crash, but LGM is the only one where I would keep an explicit fallback plan in mind if the first patch is not enough. Hugging Face https://huggingface.co/spaces/stabilityai/TripoSR/discussions/12 The current app imports rembg at module import time, and the current requirements.txt includes bare rembg but not onnxruntime . The public runtime traceback for this Space shows exactly that failure path: import rembg → import onnxruntime as ort → ModuleNotFoundError: No module named 'onnxruntime' . The README already pins python version: 3.10.13 , so Python drift is not the first thing to fix here. Hugging Face https://huggingface.co/spaces/stabilityai/TripoSR/blob/main/app.py requirements.txt omegaconf==2.3.0 Pillow==10.1.0 einops==0.7.0 transformers==4.35.0 trimesh==4.0.5 rembg +onnxruntime huggingface-hub gradio Because the failure is deterministic at startup. The current repo asks Python to import rembg before the app can even finish importing, and the public crash shows that the installed environment does not contain onnxruntime . A platform restart may have exposed it, but a clean cold start will keep hitting the same line until onnxruntime is present. This is why I would not start by upgrading Gradio or Torch here. The smallest repair is to add the missing package that the current code path actually imports. Hugging Face https://huggingface.co/spaces/stabilityai/TripoSR/blob/main/app.py not making a bigger first patch You could switch to a newer rembg extra layout, but that is not the smallest safe move for this repo. The exposed failure is not “wrong Gradio API,” not “wrong Torch version,” and not “wrong Python version.” It is specifically “ onnxruntime is missing.” So the one-line fix above is the cleanest first pass. Hugging Face https://huggingface.co/spaces/stabilityai/TripoSR/discussions/12 This Space has the same primary failure as TripoSR. app.py imports rembg , and the preprocessing path creates a rembg session. requirements.txt still lists bare rembg , and the public runtime traceback again shows ModuleNotFoundError: No module named 'onnxruntime' . Unlike TripoSR, its README metadata does not currently specify python version , even though HF supports pinning it in README YAML. Hugging Face https://huggingface.co/spaces/TencentARC/InstantMesh/blob/main/app.py README.md title: InstantMesh emoji: colorFrom: indigo colorTo: green sdk: gradio sdk version: 4.26.0 +python version: 3.10.13 app file: app.py pinned: false short description: Create a 3D model from an image in 10 seconds license: apache-2.0 requirements.txt torch==2.1.0 torchvision==0.16.0 torchaudio==2.1.0 pytorch-lightning==2.1.2 einops omegaconf deepspeed torchmetrics webdataset accelerate tensorboard PyMCubes trimesh rembg +onnxruntime transformers==4.34.1 diffusers==0.19.3 bitsandbytes imageio ffmpeg xatlas plyfile xformers==0.0.22.post7 git+https://github.com/NVlabs/nvdiffrast/ huggingface-hub Again, because the current startup path already contains the failure. The app imports rembg before the UI is ready, and the publicly reported runtime failure is the missing onnxruntime import. The Python pin is a separate hardening step: HF lets Spaces pin python version , and current ZeroGPU docs explicitly list 3.10.13 as supported. Even if the platform restart is what made the breakage visible, keeping Python fixed removes one more moving part from future cold starts. Hugging Face https://huggingface.co/spaces/TencentARC/InstantMesh/blob/main/app.py not do first I would not begin by mass-upgrading the whole dependency stack. There is already a community PR that proposes a larger cleanup including numpy<2.0.0 , Pillow==10.4.0 , newer gradio , and simplified requirements. That may be useful later, but the smallest justified first repair is still “add onnxruntime and pin Python.” Hugging Face https://huggingface.co/spaces/TencentARC/InstantMesh/discussions/32/files The next smallest hardening step is: +numpy<2.0.0 +Pillow==10.4.0 I would only do that after confirming that the startup blocker moved past rembg / onnxruntime . The reason is simple: fix the deterministic boot failure first, then deal with second-order runtime drift. Hugging Face https://huggingface.co/spaces/TencentARC/InstantMesh/discussions/32/files The current public runtime error is very specific. The Space tries to build a DDIMScheduler from stabilityai/stable-diffusion-2-1-base , and that repo id no longer resolves publicly for the needed scheduler config. The crash trace points into model/crm/model.py at the scheduler initialization. At the same time, app.py defaults --device to "cuda" and moves the model there during startup, which is an additional fragility point once the scheduler problem is fixed. Hugging Face https://huggingface.co/spaces/Zhengyi/CRM model/crm/model.py -self.scheduler = DDIMScheduler.from pretrained - "stabilityai/stable-diffusion-2-1-base", - subfolder="scheduler", - +self.scheduler = DDIMScheduler.from pretrained + "sd2-community/stable-diffusion-2-1-base", + subfolder="scheduler", + Because the current repo points at a model id that no longer works for this code path, and the public crash trace shows exactly that path failing. sd2-community/stable-diffusion-2-1-base exists, and its repo contains scheduler/scheduler config.json , which is the file CRM is trying to load. So this is not a speculative change. It is a direct one-line replacement for the dead dependency that the current startup path is trying to read. Even if a platform restart is what surfaced the error, any future cold start will keep failing until the repo id is replaced. Hugging Face https://huggingface.co/spaces/Zhengyi/CRM app.py -parser.add argument "--device", type=str, default="cuda" +parser.add argument "--device", type=str, default="cuda" if torch.cuda.is available else "cpu" The scheduler fix is the primary repair. But once the app gets past that point, startup still does model = model.to args.device and passes device=args.device into the pipeline constructor. Right now that default is hard-coded to "cuda" . So if the Space is restarted on a CPU-backed environment, or on a GPU path that is temporarily unavailable, the next boot can fail later in startup. That one-line default makes the app more robust without changing its interface or behavior when CUDA is actually available. Hugging Face https://huggingface.co/spaces/Zhengyi/CRM/blob/main/app.py not do first I would not start by adding tokens or auth logic. The current public problem is not “this repo is gated but otherwise correct.” The practical issue is that the code points at a repo id that no longer works for the scheduler path, and a community mirror already exposes the file CRM needs. So the smallest valid fix is to swap the source, not to add authentication plumbing. Hugging Face https://huggingface.co/spaces/Zhengyi/CRM LGM is the outlier. The public runtime error is not a missing Python dependency. It is a compiled-extension failure: the Space downloads its checkpoint, installs a local wheel named diff gaussian rasterization-0.0.0-cp310-cp310-linux x86 64.whl , and then crashes importing that extension because libcudart.so.11.0 is missing. The README already pins python version: 3.10.13 , so Python drift is not the first issue here. The current code also initializes most of the heavy model stack at startup, not lazily. Hugging Face https://huggingface.co/spaces/ashawkey/LGM requirements.txt torch==2.4.0 xformers numpy tyro diffusers dearpygui einops accelerate gradio imageio imageio-ffmpeg lpips matplotlib packaging Pillow pygltflib rembg gpu,cli +nvidia-cuda-runtime-cu11 rich safetensors scikit-image scikit-learn scipy tqdm transformers trimesh kiui = 0.2.3 xatlas roma plyfile app.py Add this before from core.models import LGM : python +import ctypes +import site + +for sp in site.getsitepackages : + cudart = os.path.join sp, "nvidia", "cuda runtime", "lib", "libcudart.so.11.0" + if os.path.exists cudart : + ctypes.CDLL cudart + break Because the current public crash is already precise: the installed compiled extension cannot find libcudart.so.11.0 . NVIDIA publishes nvidia-cuda-runtime-cu11 on PyPI as “CUDA Runtime native Libraries,” and this patch preloads the exact library the extension says it is missing before the extension import happens. That is the smallest repo-side change that directly matches the currently exposed failure. A platform-side restart may have exposed it, but once the process restarts, the same binary import will keep failing until the CUDA runtime library problem is addressed. Hugging Face https://huggingface.co/spaces/ashawkey/LGM This is the only one of the four where I would not promise the first patch is enough. It is the smallest targeted fix for the current public error, but native wheels can fail for more than one reason. If the wheel was built against a runtime/ABI combination that still does not match the current Spaces environment, then the next repair is no longer a one-liner. At that point, the smallest real fix becomes either: - rebuild that extension for the current runtime, or - move the Space to Docker so CUDA and the extension are under your control. HF’s current ZeroGPU docs also make clear that ZeroGPU is its own environment with H200-backed shared GPU slices and specific supported versions, so binary assumptions that worked on an older setup can stop being valid after a cold restart. Hugging Face https://huggingface.co/spaces/ashawkey/LGM not do first I would not start by upgrading Gradio, Torch, or the whole app stack just to chase this one error. The current public failure happens before any of that becomes the main issue: it dies when the compiled rasterizer tries to load C and cannot find libcudart.so.11.0 . Solve the explicit binary import error first. Then, if it boots and another error appears, fix that next one. Hugging Face https://huggingface.co/spaces/ashawkey/LGM If I were patching these repos in the smallest reasonable way, I would do exactly this: + onnxruntime README.md: + python version: 3.10.13 requirements.txt: + onnxruntime - "stabilityai/stable-diffusion-2-1-base" + "sd2-community/stable-diffusion-2-1-base" Optional second line: - default="cuda" + default="cuda" if torch.cuda.is available else "cpu" requirements.txt: + nvidia-cuda-runtime-cu11 and preload libcudart.so.11.0 before importing core.models . Hugging Face https://huggingface.co/spaces/stabilityai/TripoSR/discussions/12 Because they match the actual currently exposed startup failures , not a guessed historical failure, and because they keep the diffs small: - TripoSR: missing Python package. - InstantMesh: same missing package, plus missing Python pin. - CRM: dead external repo id. - LGM: missing CUDA runtime for a compiled extension. Hugging Face https://huggingface.co/spaces/stabilityai/TripoSR/discussions/12