arXiv:2609.35833v1 Announce Type: new Abstract: Running a language model on edge hardware provides private and low-latency reasoning without a network connection, and yet the small models that fit on such devices are unreliable on the tasks computers are expected to handle well, such as arithmetic, algebra, and formal logic problems. We argue that much of this unreliability is avoidable. Many queries appearing to demand reasoning are in fact structurally deterministic and permit fast and exact symbolic solutions. Therefore, forcing a probabilistic model to approximate them sacrifices accuracy and energy for little benefit. We present a neurosymbolic router that classifies each incoming query and dispatches it to the cheapest correct solver, sending structured tasks to deterministic engines and reserving the small language model (SLM) for open-ended word problems. Instead of hand-coding the routing logic, we learn a deterministic finite automaton (DFA) with the L* grammatical inference algorithm, using the SLM as a membership oracle and labeled data as an equivalence oracle. On a Raspberry Pi 4B (8 GB RAM, no GPU), evaluated on 100 untested prompts from DeepMind Mathematics, GSM8K, and RuleTaker, learned routing attains 100% routing accuracy and 98.3% overall accuracy with a 512-token reasoning budget (93.3% on word problems), compared with 72.0% for the strongest agent baseline, Program-of-Thought, and 58.7% for a tool-calling agent given the same solvers. Since formatted queries never reach the model, the router answers them in 1-11 ms and, in its 30-token configuration, runs 8.8x faster and 2.8x more energy-efficient than Program-of-Thought.
Neurosymbolic Routing for Reliable Reasoning on Resource-Constrained Edge Devices
A neurosymbolic router that learns a deterministic finite automaton via the L* grammatical inference algorithm achieved 100% routing accuracy and 98.3% overall accuracy on 100 untested prompts from DeepMind Mathematics, GSM8K, and RuleTaker when run on a Raspberry Pi 4B with 8 GB RAM and no GPU, according to an arXiv paper (arXiv:2609.35833v1). The router dispatches structurally deterministic queries to exact symbolic solvers and reserves the small language model for open-ended word problems, reaching 93.3% accuracy on word problems with a 512-token reasoning budget versus 72.0% for the strongest agent baseline, Program-of-Thought, and 58.7% for a tool-calling agent given the same solvers. In its 30-token configuration the router answers formatted queries in 1-11 ms and runs 8.8x faster and 2.8x more energy-efficient than Program-of-Thought.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.