cd /news/ai-infrastructure/moorcheh-edge-on-arduino-uno-q-vs-ve… · home › topics › ai-infrastructure › article
[ARTICLE · art-144104] src=moorcheh.ai ↗ pub= topic=ai-infrastructure verified=true sentiment=· neutral

Moorcheh Edge on Arduino UNO Q vs. Ventuno Q: 1536-D Vector Benchmark

Moorcheh Edge's 1536-dimensional vector benchmark on two ARM64 edge boards found that at 1 million vectors the Ventuno Q (about 15 GB RAM) was roughly 3× faster on search (269 ms vs 833 ms median) and about 4× faster on full ingest (8 hours vs 32 hours wall clock) than the Arduino UNO Q (about 3.6 GB RAM), while at 10k vectors median search latency tied at about 25 ms on both. The benchmark ran Moorcheh Edge in vector-only mode with MIB binarization at ingest and EDM retrieval, keeping the 1M-vector store at about 319 MiB on disk and about 1.2–1.5 GiB in Docker while serving search. Both boards received identical software and test parameters (top_k = 5, seed 42, 5 warmup plus 50 timed searches), with the store limit raised to 1,000,000 items for the stress test.

by read7 min views1 publishedOct 2, 2026
Moorcheh Edge on Arduino UNO Q vs. Ventuno Q: 1536-D Vector Benchmark
Image: source

← Back to Blog We compared two ARM64 edge boards running the same Moorcheh Edge workload: Arduino UNO Q (about 3.6 GB RAM) and Ventuno Q (about 15 GB RAM). Each board ingested and searched 1536-dimensional vectors in vector mode only — no text pipeline, no embedding model, no LLM. Queries were raw float vectors sent to POST /search.

For this benchmark we raised the per-store item limit to 1,000,000 so we could stress-test ingest and search at scale. Typical Moorcheh Edge deployments use a smaller cap suited to kiosk-sized catalogs. Even with Moorcheh's compact storage, 1M vectors is an extreme case; it shows how the two boards behave when the catalog grows.

For the earlier UNO Q RAG pipeline breakdown (embed + LLM vs retrieval), see [Edge RAG on 4 GB Hardware](https://moorcheh.ai/blog/edge-rag-4gb-hardware-arduino-uno-q-benchmarks).

Headline results: At 10k vectors, median search latency is a tie (about 25 ms on both). At 1M, Ventuno Q is about 3× faster on search (269 ms vs 833 ms median) and about 4× faster on full ingest (8 hours vs 32 hours wall clock).

Detailed tables (upload progress, 50 search samples per run, RAM during ingest, and the 50-query search + RAM study at 1M) are in the Google Sheets linked under Full results.

Why Moorcheh Edge makes this benchmark possible #

A conventional 1M × 1536 float32 vector store is roughly 6 GB of embedding data alone (1,000,000 × 1536 × 4 bytes), before index structures and runtime overhead. That does not fit the practical memory budget on Arduino UNO Q and is a poor use of Ventuno Q if the goal is edge deployment rather than raw storage.

Moorcheh Edge binarizes vectors at ingest with MIB (Moorcheh Information Binarization) and retrieves with EDM (Enhanced Distance Metric) over those compact forms. In our runs:

  • On disk: about 3.2 MiB (10k), about 32 MiB (100k), about 319 MiB (1M)
  • In Docker at 1M: about 1.2–1.5 GiB while serving search

The same item counts on both boards confirm that store size is driven by Moorcheh's encoding, not by which CPU runs the server. That small footprint is why we could run a 1M × 1536 stress test on real edge hardware instead of only on a cloud instance.

How we ran the benchmark #

We used one Moorcheh Edge server deployment on each board (ARM64, vector-only stack). The client ran on the board against localhost, so timings include HTTP JSON plus server work.

Procedure for each catalog size (10k, 100k, 1M):

  1. Clear the store (fresh run at that scale).
  2. Upload random unit vectors in batches until the target count is reached; record wall time and throughput.
  3. Sample host and container memory during ingest.
  4. Run 5 warmup searches, then 50 timed searches with top_k = 5 , fixedseed = 42 , and report min, median, mean, p95, and p99.

After the 1M run completed on each board, we ran a separate 50-query study on the loaded 1M catalog: before and after each search, record latency and RAM (host used memory and Docker stats for the Moorcheh container).

Both boards received the same software and the same test parameters; only hardware resources differ.

Full results #

Every run is documented in two public Google Sheets — each contains the summary, per-query search latencies, upload progress, RAM samples during ingest, and the 50-query RAM worksheet:

Test configuration #

Parameter Value
Server Moorcheh Edge (same deployment on both boards)
Store limit (this benchmark) 1,000,000 items (raised for stress testing)
Store mode Vector (precomputed floats from client)
Dimension 1536 (locked on first upload)
Query Random unit vector, seed 42
Search top_k = 5, 5 warmup + 50 timed runs
Catalog sizes tested 10,000 / 100,000 / 1,000,000 (separate runs)
Device RAM context (from measurements)
Arduino UNO Q Docker limit about 3.58 GiB; strong pressure at 1M
Ventuno Q Docker limit about 14.93 GiB; same Moorcheh server, more headroom

Results: store size #

On-disk store size was identical on both boards at every scale.

On-disk store size: Moorcheh MIB store vs raw float32 (MiB)

Green: Moorcheh store (identical on both boards). Red: raw float32 embeddings at 1M × 1536 — ~19× larger than the Moorcheh store, before index structures or runtime overhead.

Moorcheh Edge loads the binarized catalog into memory for search; it does not keep full float32 matrices for every vector.

Results: upload (ingest) #

Vectors Arduino UNO Q Ventuno Q Ventuno advantage
10k 59.3 s 29.5 s about 2.0× faster
100k about 19 min about 10 min about 1.9× faster
1M about 32.1 h about 8.0 h about 4.0× faster

At 10k, UNO Q ingest stays practical (about 1 minute). At 1M, ingest dominates wall clock: UNO Q needs on the order of 32 hours for a full load; Ventuno about 8 hours.

Peak host memory during ingest:

| Scale | UNO Q peak (MB) | Ventuno peak (MB) | 
|---|---|---|

| 10k | 817 | 1,028 | | 100k | 875 | 1,155 | | 1M | 2,673 | 2,708 |

Both boards reach about 2.7 GB host use at 1M. UNO Q is near its ceiling; Ventuno still has room below its Docker limit.

Results: search latency (50 runs per scale) #

Median search latency by catalog size (ms, top_k = 5, 50 runs)

Tie at 10k; at 1M Ventuno Q is ~3.1× faster (269 vs 833 ms). p99 and per-run samples are in the linked spreadsheets.

Search cost grows with N. At kiosk scale (10k), both boards deliver about 25 ms median, which is suitable for interactive edge retrieval. At 1M, UNO Q is about 0.83 s per query vs Ventuno about 0.27 s. Latency scales roughly linearly with catalog size, consistent with scanning the full in-memory binarized index.

With 1M vectors already loaded, we measured each of 50 searches together with RAM immediately before and after the request. Typical host RAM stayed around 2.26 GB on UNO Q and 2.11 GB on Ventuno Q.

50-query study on the loaded 1M catalog

Median latency (ms)

Docker memory (% of container limit)

Similar index footprint (~1.2–1.5 GiB) on both boards — Ventuno's extra RAM is headroom, not a larger index. All 50 rows are in the linked spreadsheets.

Per-query RAM barely changes during search because the catalog is already resident. The gap is CPU and system throughput, not index size.

Choosing a board #

Arduino UNO Q fits Moorcheh Edge at 10k (and similar) catalog sizes: about 25 ms search, about 1 min ingest, about 3 MiB store. 100k is usable if about 100 ms retrieval is acceptable. 1M is viable as a stress experiment on this hardware but not as a default production target without long ingest times and about 800 ms-class search.

Ventuno Q runs the same Moorcheh Edge stack but is the better match for 100k–1M catalogs: lower search latency at scale and much shorter 1M ingest.

Methodology and limitations #

  1. Search timings include localhost HTTP JSON, not internal Rust-only timers.
  2. Each run uses the same seed and query vector for repeatability, not a mix of query types.
  3. We report vector ingest, search speed, and memory only (no text RAG quality or LLM tests).
  4. Ventuno Q is our name for a higher-RAM ARM64 edge peer in this study.
  5. Product store limits for typical edge use remain below the 1M cap used here; we raised the limit only to stress-test Moorcheh Edge at extreme scale.

Try Moorcheh Edge #

- [On-Edge introduction](https://docs.moorcheh.ai/on-edge/introduction)
- [Edge product overview](https://moorcheh.ai/products/edge)
- Research: [From HNSW to Information-Theoretic Binarization](https://arxiv.org/pdf/2601.11557)

Benchmarks run 29 September – 1 October 2026.

Build this architecture today.

Get your API key and start building agentic memory in under 5 minutes.

Get API Key

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @moorcheh edge 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/moorcheh-edge-on-ard…] indexed:0 read:7min 2026-10-02 · —