Building Darb: A Dual-Engine AI Architecture for Real-Time Assistive Vision Software engineer Abdulrahman Haramain built Darb, an assistive navigation platform for the visually impaired that splits inference between a zero-latency on-device vision engine running MediaPipe at 30+ FPS and a cloud-based semantic engine on Vertex AI that returns rich scene descriptions in one to three seconds. Custom state-management hooks let the fast reflex engine interrupt and override the slower cloud comprehension engine when a hazard is detected, so users get immediate obstacle warnings without losing contextual awareness. The project won an Outstanding Completion Award at Kuwait Codes 2026. Have you ever tried walking around your room with your eyes closed? Now imagine relying on a cloud AI model to guide you. If it takes 2 seconds to process a video frame and tell you there's a wall ahead... you've already walked into it. When it comes to assistive tech, latency isn't just an annoyance—it’s a safety hazard. But if you rely only on offline, on-device models, you lose out on the insane contextual awareness that modern LLMs give you. I ran straight into this problem while building Darb, an assistive navigation platform for the visually impaired which ended up taking home an Outstanding Completion Award at Kuwait Codes 2026 . My solution? Stop making one AI do everything. I split Darb's "brain" into two engines. "Reflexes" vs. "Comprehension" Think about human biology. If someone throws a ball at your face, you duck instantly a reflex . You don't consciously analyze the ball's trajectory first. But when you walk into a new coffee shop, you take a second to scan the room and figure out where the empty tables are comprehension . Darb mimics this using two separate inference engines: The Job: Catch immediate obstacles, barriers, or drops. The Speed: 30+ FPS. Zero latency. The Best Part: It works completely offline. No waiting on network requests when you're about to trip. The Job: Semantic scene understanding. The Speed: ~1 to 3 seconds depending on the network . The Output: Rich audio descriptions like, "You are in a hallway. There is an open door on your right, and a person walking towards you." The Stack That Makes It Work Core: React and Next.js fast, snappy, and scalable . Vision: MediaPipe for the zero-latency client-side spatial mapping . Cloud: Vertex AI for the heavy-lifting semantic analysis . The Glue: I wrote custom state-management hooks so the "Reflex" engine can instantly interrupt and override the "Comprehension" engine if it spots a hazard while the cloud is still "thinking." The Biggest Lesson Building Darb taught me that AI engineering is way more than just pinging the smartest API you can find. It’s about actual systems architecture. It’s about knowing where the computing should happen based on the real-world physical constraints of the person using your app. By mixing the raw speed of Edge ML with the depth of Vertex AI, I didn't have to choose between keeping users safe and keeping them informed. Have you guys messed around with hybrid edge/cloud ML setups? I'd love to hear how you handle the latency vs. context trade-off Abdulrahman Haramain is a software engineer and AI systems developer based in Kuwait. He builds production web platforms and offline computer-vision applications. Connect on GitHub, LinkedIn, or at abdulrahmanharamain.com.