{"slug": "mlx-smolvla-smolvla-checkpoint-execution-and-validation-on-apple-silicon", "title": "Mlx-smolvla: SmolVLA checkpoint execution and validation on Apple Silicon", "summary": "A metadata fixture test on Linux/CPU found that the mlx-smolvla code path's fixed-filename processor state lookup treats a published SmolVLA fine-tune's step-6 normalizer as identity instead of mean_std, while lookup via the serialized state_file returns mean_std. The checkpoint in question, jackvial/so101_smolvla_pickplaceorangecube_e100, declares its normalizer state as policy_preprocessor_step_6_normalizer_processor.safetensors with STATE/ACTION using MEAN_STD, whereas mlx-smolvla assumes the canonical step-5 preprocessor normalizer and step-0 postprocessor unnormalizer layout. The author states this is not a Metal reproduction or a proven full-checkpoint loading failure, but a silent compatibility edge, and recommends making the serialized processor contract authoritative for current-format checkpoints.", "body_md": "Hmm… I don’t have a Mac, so I can’t test Metal directly, but for now:\n\nOn the validation split: yes, the **strict CPU / production Metal separation makes sense to me**, as long as the guarantees of the two lanes stay explicit and separate.\n\nI would think of it as a small validation ladder rather than one global definition of “same”:\n\n1. the serialized checkpoint/processor contract is coherent;\n2. MLX CPU matches the PyTorch reference under controlled inputs/noise;\n3. Metal’s direct numerical divergence is measured separately;\n4. an offline functional/non-regression metric checks whether that divergence actually degrades the policy output;\n5. eventually, closed-loop robot behavior is a separate layer again.\n\nThat also avoids making bitwise CPU↔accelerator equality the production requirement. PyTorch makes the same general warning in its [numerical accuracy notes](https://docs.pytorch.org/docs/main/notes/numerical_accuracy.html): mathematically equivalent floating-point computations are not guaranteed to be bitwise identical across devices/backends.\n\nSo I would keep the strict negative Metal results rather than “fixing” them by widening the tolerance. A failed strict gate is still useful evidence; the production lane can answer a different question.\n\nThere is also one concrete fine-tuned-checkpoint edge that may be useful for the other question you asked.\n\n### \n\nI found a public SmolVLA fine-tune, [`jackvial/so101_smolvla_pickplaceorangecube_e100`](https://huggingface.co/jackvial/so101_smolvla_pickplaceorangecube_e100/commit/c7bc01d727947811e7b708e8aabe613da76ab4e9), whose preprocessor contains an additional step before normalization.\n\nIts `policy_preprocessor.json` therefore declares the normalizer state as:\n\n`policy_preprocessor_step_6_normalizer_processor.safetensors`\n\nand that file actually exists. The processor configuration also says STATE/ACTION use `MEAN_STD`.\n\nThis matters because LeRobot does not appear to treat “normalizer == step 5” as the semantic contract. In the current [`processor/factory.py`](https://github.com/huggingface/lerobot/blob/main/src/lerobot/processor/factory.py), policies explicitly compose processor steps in their own order, and the source describes that **step order as a Hub-serialized contract**. Checkpoints likewise contain `policy_preprocessor_step_*.safetensors` / `policy_postprocessor_step_*.safetensors`, rather than promising one universal normalizer index; see the current [checkpoint layout docs](https://github.com/huggingface/lerobot/blob/main/docs/source/multi_gpu_training.mdx).\n\nThe serialized processor JSON already carries the actual `state_file`, and LeRobot’s [` processor/pipeline.py`](https://github.com/huggingface/lerobot/blob/main/src/lerobot/processor/pipeline.py) uses that serialized state information when reconstructing the pipeline.\n\nBy contrast, in the `mlx-smolvla` code path I checked, processor state discovery is still based on the canonical fixed filenames around:\n\nwith the usual step-5 preprocessor normalizer / step-0 postprocessor unnormalizer layout.\n\nI tested this part on Linux/CPU only, using a metadata fixture matching the published step-6 layout. The result was:\n\n| layout | current fixed lookup | lookup via serialized `state_file` | \n| canonical step-5 normalizer | `mean_std` | `mean_std` | \n| published-style step-6 normalizer | `identity` | `mean_std` | \n\n So this is **not a Metal repro**, and I also would not call it a proven full-checkpoint loading failure yet. The narrower finding is:\n\na checkpoint can look acceptable at the core-config level while a valid serialized state normalizer lives at a different processor step, causing the current fixed filename lookup to treat state normalization as identity.\n\nThat looks like a useful compatibility edge because it can be silent rather than producing an obvious load exception.\n\nMy default route here would be to make the serialized processor contract authoritative for current-format checkpoints:\n\n- read `policy_preprocessor.json` /`policy_postprocessor.json` ;\n- identify the relevant `normalizer_processor` /`unnormalizer_processor` step;\n- follow that step’s declared `state_file` ;\n- include that declared file in the Hub download set;\n- then keep the existing tensor/header/feature-key checks after resolving the file.\n\nI would not just change the magic number from step 5 to step 6, since the whole point is that the index depends on the composed processor pipeline.\n\nA useful failure split might be:\n\n- **processor JSON + declared state file present** → use the declared file;\n- **legacy checkpoint without processor JSON** → use a clearly separate legacy/migration path;\n- **processor JSON declares state, but that state file is missing** → explicit diagnostic rather than guessing;\n- **state file is present, but feature/stat keys do not match** → report that separately as a schema/stats mismatch.\n\nThat last distinction may be worth preserving because LeRobot itself currently warns that processor/normalization mismatches can run without raising an error while silently damaging results; see [Adding a Policy](https://huggingface.co/docs/lerobot/bring_your_own_policies). There is also a recent concrete example in [LeRobot #4415](https://github.com/huggingface/lerobot/issues/4415), where normalization/unnormalization can be silently skipped because the serialized stats keys do not match the runtime feature key.\n\nSo for “compatible checkpoint” reports, I think the processor artifacts are part of the compatibility contract, not just ancillary files next to `model.safetensors`.\n\nA compact triage order might therefore be:\n\n1. architecture/config supported;\n2. required weight tensors/names/shapes supported;\n3. processor JSON and its declared state files available;\n4. camera/state/action schema and rename mapping compatible;\n5. normalization mode + stats keys/shapes compatible;\n6. only then investigate MLX CPU/Metal numerical behavior.\n\nThat should make future “this fine-tune doesn’t work” reports cheaper to classify.\n\n## \nSmall public-checkpoint scan\n\nI also did a metadata-only CPU scan across a convenience sample of public Hub repositories; no model weights or Apple execution were involved.\n\nThe scan saw:\n\n- 86 candidate repositories;\n- 72 with `config.json` identifying`type == smolvla` ;\n- 68 accepted by the dependency-light `mlx-smolvla` config parser;\n- 0 scan errors;\n- 1 strong case where the core config was accepted, an active serialized normalizer state existed at another step, and the fixed lookup resolved state normalization to `identity` .\n\nThat strong case was the `jackvial/...` step-6 checkpoint above.\n\nI would **not** interpret `1/86` as a prevalence estimate: the Hub search was not a census, and many of the candidate repositories were outside the native port’s full compatibility surface for unrelated reasons. The useful result is just that the edge exists in a real published SmolVLA checkpoint rather than only as a synthetic possibility.\n\n \n### \n\nI would keep this lane, but I think its meaning is slightly different from direct backend parity.\n\nAs I read [`scripts/statistical_check.py`](https://github.com/daniiarabdiev/mlx-smolvla/blob/main/scripts/statistical_check.py), it compares each backend’s first predicted action with the dataset action and then compares the resulting MAEs.\n\nSo approximately it answers:\n\n“Does production MLX make this offline imitation-error proxy worse than the PyTorch reference?”\n\nThat is useful, but it is distinct from:\n\n“How far did the MLX prediction itself move from the PyTorch prediction?”\n\nI think keeping both concepts separate actually strengthens the validation story.\n\nA cheap addition, if useful, would be to reuse the same deterministic 50-frame inputs/noise and record a direct MLX↔Torch action-difference distribution too — for example median / p95 / max absolute delta, perhaps per action dimension or over the whole action chunk.\n\nI would not replace the existing 8-case direct gate or invent a new pass threshold from that distribution. It would mainly make the distinction visible:\n\n- **direct numerical disagreement** ;\n- **dataset-relative offline non-regression** .\n\nThen if Metal differs numerically but the offline metric remains stable, that is a more informative result than either number alone.\n\n## \nWhy I would keep strict numerical and downstream checks separate on M5\n\nThere are some recent MLX examples that make this distinction useful, although I would not assume they have the same root cause as `mlx-smolvla`.\n\nIn [MLX #3897](https://github.com/ml-explore/mlx/issues/3897), batched/masked attention on M5 differs enough from the single-sequence path to fail a strict equivalence test, while the reported selected token remains unchanged. The same discrepancy is much smaller on M3 Max.\n\nThat is a good example of:\n\n**strict numerical mismatch → observed downstream choice still unchanged**\n\nBut the opposite case exists too. [MLX #3953](https://github.com/ml-explore/mlx/issues/3953) isolated a real float32 correctness problem in a broadcasted matmul path. A very simple invariant made the distinction possible: in the reported `L=1` attention case, the output mathematically has to equal `V`, but it did not.\n\nSo “strict test failed” alone does not tell us whether a difference is harmless rounding, a different-but-acceptable backend path, or an actual correctness bug.\n\nThat is why the current separation into a reference lane and a production-behavior lane seems useful to me.\n\nIf you ever want one very cheap M5-side discriminator, `MLX_ENABLE_TF32=0` might also be worth a single A/B run using the *same* fixed inputs/noise.\n\nThis is only a diagnostic hypothesis, not a proposed root cause. A different MLX project reported M5 equivalence tests that failed normally but passed with `MLX_ENABLE_TF32=0` in [mlx-swift-lm #357](https://github.com/ml-explore/mlx-swift-lm/issues/357). MLX has also had discussion around the flag and fp32 precision behavior, e.g. [MLX #3860](https://github.com/ml-explore/mlx/issues/3860).\n\nSo I would use it only like this:\n\n- default environment → record the existing Vision/action deltas + latency;\n- `MLX_ENABLE_TF32=0` before MLX initialization → repeat exactly;\n- if the delta changes substantially, that narrows the kernel/precision-path investigation;\n- if it does not, that hypothesis loses weight.\n\nI would not infer from another project’s result that the current SmolVLA Metal delta is “a TF32 bug”.\n\n \n## \nWhere offline parity eventually stops\n\nOne final boundary that I think your current wording already handles sensibly: offline action parity is not yet sustained robot task success.\n\nThat distinction matters especially for SmolVLA because deployment is chunked/closed-loop rather than a sequence of independent one-step predictions.\n\nLeRobot’s current [asynchronous inference documentation](https://huggingface.co/docs/lerobot/async) makes this explicit: `actions_per_chunk`, `chunk_size_threshold`, overlap aggregation, inference latency, and the state of the action queue can all change behavior. It even notes that larger action horizons can accumulate prediction error.\n\nSo if the project eventually wants a stronger production claim, I would see the next validation layer as paired/repeated rollout behavior under controlled initial conditions and the actual serving mode — not simply tightening the numerical tolerance further.\n\nI would not make that a requirement for the current milestone; it is just the natural endpoint of the validation ladder.\n\n \nSo if I had to pick the lowest-cost next steps, I would probably keep the current strict/production split, preserve the existing 50-frame offline metric, optionally add a direct 50-frame MLX↔Torch delta summary, and make processor-state resolution follow the serialized `state_file` for current-format LeRobot checkpoints.\n\nThe processor-state case seems especially useful because it is independent of Apple hardware and gives a concrete compatibility condition that future checkpoint reports can test.", "url": "https://wpnews.pro/news/mlx-smolvla-smolvla-checkpoint-execution-and-validation-on-apple-silicon", "canonical_source": "https://discuss.huggingface.co/t/mlx-smolvla-smolvla-checkpoint-execution-and-validation-on-apple-silicon/179927#post_4", "published_at": "2026-10-04 19:11:10+00:00", "updated_at": "2026-10-04 19:41:23.391275+00:00", "lang": "en", "topics": ["robotics", "machine-learning", "ai-tools"], "entities": ["mlx-smolvla", "SmolVLA", "Apple Silicon", "Metal", "PyTorch", "LeRobot", "Hugging Face", "jackvial/so101_smolvla_pickplaceorangecube_e100"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/mlx-smolvla-smolvla-checkpoint-execution-and-validation-on-apple-silicon", "markdown": "https://wpnews.pro/news/mlx-smolvla-smolvla-checkpoint-execution-and-validation-on-apple-silicon.md", "text": "https://wpnews.pro/news/mlx-smolvla-smolvla-checkpoint-execution-and-validation-on-apple-silicon.txt", "jsonld": "https://wpnews.pro/news/mlx-smolvla-smolvla-checkpoint-execution-and-validation-on-apple-silicon.jsonld"}}