{"slug": "show-hn-sovereign-nb-sub-microsecond-bare-metal-neural-engine-in-rust-uefi", "title": "Show HN: Sovereign-NB: Sub-microsecond bare-metal neural engine in Rust (UEFI)", "summary": "A developer released Sovereign-NB, a bare-metal x86-64 neural inference appliance written in Rust that boots as a UEFI application and serves fixed-size input frames over a UART stream. The runtime runs two integer-only model paths — a 64→32→16 ternary MLP and 16-dimensional causal linear attention with a persistent 16×16 recurrent state — from a no_std, no-heap core built for the x86_64-unknown-uefi target with a nightly toolchain. Models are packed as NEUR shards with a 16-byte header and two-bit ternary weights (640 bytes for the MLP, 832 bytes for attention), and can be swapped by replacing NEURAL_DATA:\\weights.bin on a FAT32 volume before the next boot.", "body_md": "Sovereign Neural Box is a small, bare-metal x86-64 inference appliance. It boots as a UEFI application, loads a packed ternary model, and serves fixed-size input frames over a UART stream. The runtime combines two integer-only model paths: a feed-forward MLP and recurrent causal linear attention.\n\n- **Dual-architecture inference:** 64→32→16 ternary MLP and 16-dimensional causal linear attention with a persistent 16×16 recurrent state.\n- **No-heap inference core:**`src/main.rs` is`#![no_std]` ; the core does not enable Rust`alloc` or use`Vec` /heap allocation. Model, DMA, frame, and state storage use fixed-size buffers. UEFI file I/O writes directly into the preallocated DMA-aligned shard buffer.\n- **UEFI x86-64 target:** built for`x86_64-unknown-uefi` with the repository's nightly toolchain.\n- **User-friendly model updates:** the GPT appliance image has a FAT32 EFI System Partition and a FAT32`NEURAL_DATA` volume. Replace`NEURAL_DATA:\\weights.bin` to load a different model at the next boot.\n- **Fallback behavior:** UEFI SimpleFileSystem volumes are searched before`ExitBootServices` ; invalid or missing files fall back to legacy raw-NVMe shard lookup, then to a safe built-in identity model.\n- **UART streaming:** COM2 accepts 64 signed-byte inputs and returns 16 little-endian`i32` outputs. The standalone`NR` control marker resets recurrent attention state.\n\n- `src/` — UEFI entry point, model kernels, shard parsing, UART, NVMe, and shared-memory support.\n- `tools/package_image.py` — GPT/FAT32 appliance image builder;`tools/package_image.ps1` is its PowerShell wrapper.\n- `tools/payload_builder/` — host-side Rust NEUR shard generator (`mlp` or`attention` ).\n- `tools/test_dual_volume.ps1` — QEMU test for FAT-based model loading and UART streaming.\n- `tools/test_attention_sequence.ps1` — QEMU test for recurrent attention accumulation and reset.\n- `ml/` — optional model training and export utilities.\n- `DEPLOYMENT.md` — detailed flashing, model-update, and server deployment guidance.\n\nA NEUR shard begins with a 16-byte header: ASCII magic `NEUR`, little-endian version and input dimension, a model-type byte, little-endian output dimension, and a final hidden/attention dimension byte. Ternary weights use two bits per weight: `00` is zero, `01` is +1, and `11` is −1.\n\n| Model type | Value | Dimensions | Packed payload | \n|---|---|---|---|\n| Ternary MLP | `0` | 64 → 32 → 16 | 640 bytes | \n| Causal linear attention | `1` | Q/K/V: 64 → 16; O: 16 → 16 | 832 bytes | \n\nFor attention, the recurrent state is a row-major 16×16 matrix of `i32` values. It accumulates key/value outer products across frames and is reset by the UART control marker.\n\nInstall the Rust nightly toolchain and the UEFI target listed in `rust-toolchain.toml`, then build the release EFI application:\n\n```\ncargo +nightly build --target x86_64-unknown-uefi --release\n```\n\nThe core is `no_std` and uses fixed storage for model weights, I/O frames, and attention state. File-system protocol metadata may be managed internally by UEFI firmware; no Rust heap allocator is enabled by this crate.\n\nBuild the EFI binary first. The packager uses `dist/production_shard.bin` when present, or creates a small valid default MLP shard if it is absent.\n\n```\npython tools/package_image.py\n```\n\nThis creates `dist/neural_box_appliance.img`, a GPT disk image with a protective MBR:\n\n1. **ESP:** FAT32, contains`\\EFI\\BOOT\\BOOTX64.EFI` and`\\STARTUP.NSH` .\n2. **NEURAL_DATA:** FAT32, contains the default model as`\\weights.bin` .\n\nFor an Attention shard, generate it with the host builder and pass it as the packager's shard input (the default packaging path is `dist/production_shard.bin`):\n\n```\ncargo run --manifest-path tools/payload_builder/Cargo.toml --target x86_64-pc-windows-msvc --release -- --model attention --output dist/production_shard.bin\npython tools/package_image.py\n```\n\nThe image can also be prepared with `tools/package_image.ps1`. See `DEPLOYMENT.md` before writing an image to physical media.\n\nBefore leaving Boot Services, the application scans UEFI `SimpleFileSystem` handles using a fixed caller-owned handle array and searches for `\\weights.bin` or `\\NEURAL_WEIGHTS\\weights.bin`. A valid shard is read directly into the 4 KiB DMA-aligned buffer and dispatched by its model-type field. If no file is found or parsing fails, the application attempts the legacy raw-NVMe locations; if that also fails, it runs the safe built-in fallback model.\n\nTo update a deployed appliance, mount the `NEURAL_DATA` FAT32 volume on a desktop OS and replace its root `weights.bin` with a valid NEUR shard. No EFI partition modification is needed.\n\n- **Input:**`NB` followed by exactly 64 raw signed`i8` bytes.\n- **Output:**`NR` , one dimension byte, then that many little-endian`i32` values (16 for both supported models).\n- **Attention reset:** send standalone`NR` with no following payload. This is a control event and does not produce an output frame.\n\nThe UEFI streaming loop has a finite frame limit and timeout intended for appliance/QEMU verification. COM1 carries text diagnostics; COM2 carries the binary protocol.\n\nWith QEMU installed and `assets/OVMF.fd` available:\n\n```\ncargo +nightly build --target x86_64-unknown-uefi --release\npython tools/package_image.py\n.\\tools\\test_dual_volume.ps1\n```\n\nThe dual-volume test boots the GPT image with the disk attached as NVMe, sends eight input frames, verifies 16-element responses, and asserts that COM1 reports a FAT-loaded shard and zero dropped frames. To verify attention state and reset:\n\n```\n.\\tools\\test_attention_sequence.ps1\n```\n\nLicensed under either of:\n\nat your option.", "url": "https://wpnews.pro/news/show-hn-sovereign-nb-sub-microsecond-bare-metal-neural-engine-in-rust-uefi", "canonical_source": "https://github.com/msndsn123a/Sovereign-NB", "published_at": "2026-10-02 01:16:11+00:00", "updated_at": "2026-10-02 01:46:04.332110+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-infrastructure", "developer-tools", "neural-networks"], "entities": ["Sovereign-NB", "Rust", "UEFI", "NEUR", "x86_64-unknown-uefi", "NEURAL_DATA", "QEMU"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-sovereign-nb-sub-microsecond-bare-metal-neural-engine-in-rust-uefi", "markdown": "https://wpnews.pro/news/show-hn-sovereign-nb-sub-microsecond-bare-metal-neural-engine-in-rust-uefi.md", "text": "https://wpnews.pro/news/show-hn-sovereign-nb-sub-microsecond-bare-metal-neural-engine-in-rust-uefi.txt", "jsonld": "https://wpnews.pro/news/show-hn-sovereign-nb-sub-microsecond-bare-metal-neural-engine-in-rust-uefi.jsonld"}}