{"slug": "fitguru-building-a-private-on-device-ai-workout-coach-in-kotlin-multimodel-ai", "title": "FitGuru: Building a Private, On-Device AI Workout Coach in Kotlin(Multimodel AI)", "summary": "Team ARKA built FitGuru, an Android workout coach that runs MediaPipe pose estimation entirely on-device, using joint-angle state machines for rep counting and form cues while only sending structured JSON (angles, reps, score, heart rate) off the phone for optional Gemini debriefs and ElevenLabs voice. The architecture splits into an always-on Reflex Engine for per-frame analysis and a Cognitive Engine that activates after a set, keeping camera frames and 33 pose landmarks local. The project was submitted at Innovhacks 4.0 and supports live form analysis for 18 exercises.", "body_md": "published: false\n\ndescription: \"We built a phone-side workout coach at Innovhacks 4.0 - MediaPipe pose, joint-angle state machines, Gemini debriefs, and ElevenLabs voice - without sending your camera feed to the cloud.\"\n\ntags: ai, genai, agentaichallenge, kotlin\n\n**Innovhacks 4.0 · Team ARKA · Android (Kotlin) · Multimodal AI**\n\nYou are three reps into a hard set. You cannot look at a screen. You also cannot tell if your squat is deep enough, if your chest collapsed, or if that last lockout even counted.\n\nMost “smart” fitness apps solve this by shipping your camera to a server. That is a lot to ask for video of your body in a gym.\n\nAt **Innovhacks 4.0**, Team ARKA built **FitGuru** to test a narrower idea:\n\nCan a phone watch your form, count the set, talk you through mistakes, and debrief you afterwards - **without the video ever leaving the device?**\n\nThe short answer: yes, if you split the coach into two brains. A cheap **Reflex Engine** that runs on every camera frame. And a **Cognitive Engine** that only wakes up when the set is over, and only sees numbers.\n\n**Demo:**[https://drive.google.com/file/d/1yyQeH4TjNK_Y5u2S9msBnvTxCkJeANAF/view?usp=sharing](https://drive.google.com/file/d/1yyQeH4TjNK_Y5u2S9msBnvTxCkJeANAF/view?usp=sharing)\n\n**Code:** [github.com/atharvachaudhari1/innovhacks4.0_arka](https://github.com/atharvachaudhari1/innovhacks4.0_arka)\n\n**APK:** `apk/FitGuru-Innovhacks4.apk`\n\nFitGuru was our submission to Innovhacks 4.0 through Major League Hacking (MLH). We had a weekend and one hard rule (the video never leaves the phone), which is why the architecture ended up as two brains instead of one giant model.\n\nWe also used two of the tools MLH hackathons often put in front of builders:\n\nIf the camera feed leaves the phone, we already lost.\n\nThat one rule decided the architecture:\n\n| Must stay on the phone | Allowed off the phone (opt-in) | \n|---|---|\n| Camera frames | Nothing visual | \n| 33 pose landmarks | Nothing visual | \n| Rep counting + form cues | Structured JSON of angles, reps, score, heart rate | \n| Workout history (Room) | Gemini debrief text | \n| Offline TTS fallback | ElevenLabs audio, if the user configured a key | \n\nNo account. No cloud video pipeline. If the network dies mid-session, the coach still counts reps and still talks.\n\nOpen the app, pick a lift, put the phone where it can see you, and start the set.\n\n**Live form analysis for 18 exercises**\n\nSquats, push-ups, deadlifts, pull-ups, chin-ups, lunges, plank, bicep curls, running cadence, bench / chest press, shoulder press, leg press / extension / curl, lat pulldown, seated row, triceps pushdown.\n\n**Automatic rep counting** using joint-angle state machines. A squat is not a guess. A rep exists only when the knee and hip cross a “down” threshold and then return past an “up” threshold.\n\n**Spoken corrections** so you never look at the screen mid-set. Short cues during the set. A longer spoken review after you say *“finish set.”*\n\n**Hands-free voice commands** - *finish set*, *smart rest*, *form check*, *switch to deadlift*, *weight 40*.\n\n**Local history** - reps, form score, volume, calories, average heart rate - stored in Room. Daily challenges and a weekly rep goal live on the dashboard.\n\n**A 20-machine gym guide** for beginners: what the machine does, how to set it up, common mistakes, and whether FitGuru can coach that movement live.\n\n**A web companion** with the same MediaPipe loop in the browser, plus an iOS SwiftUI client that uses Apple Vision.\n\nThe tempting hackathon move is: stream frames to an LLM and ask it “is this squat good?”\n\nThat fails in three ways at once.\n\nSo FitGuru is multimodal by **role**, not by dumping every sensor into one prompt.\n\n```\nCamera ──► MediaPipe Pose (on-device)\n                │\n                ▼\n         33 landmarks / frame\n                │\n                ▼\n        Reflex Engine (always on)\n        • joint angles\n        • state machine (ready → down → up)\n        • form faults + 2.5s cue cooldown\n        • Gym Shield (lock onto you, ignore walkers)\n                │\n     ┌──────────┴──────────┐\n     ▼                     ▼\n ElevenLabs / TTS      Room (history)\n     │\n     ▼  only after the set\n Gemini 2.5 Flash\n (JSON of numbers, never video)\n     │\n     ▼\n Spoken debrief + recovery protocol\n```\n\n**Reflex** is trigonometry and thresholds. It is cheap enough to run on every frame.\n\n**Cognitive** is Gemini. It sees a session summary, not a pixel.\n\nThat split is the product.\n\nThe Android camera path is CameraX → MediaPipe Pose Landmarker (`pose_landmarker_full.task`) in `LIVE_STREAM` mode. Each result is 33 normalized 3D landmarks: nose, shoulders, elbows, wrists, hips, knees, ankles, toes.\n\nWe never ask the model “was that a squat.” We compute the interior angle at a joint.\n\n```\n// angle at landmark B, formed by A–B–C\nfun calculateAngle(a: Landmark, b: Landmark, c: Landmark): Double {\n    val v1x = a.x() - b.x(); val v1y = a.y() - b.y()\n    val v2x = c.x() - b.x(); val v2y = c.y() - b.y()\n    val dot = v1x * v2x + v1y * v2y\n    val mag1 = hypot(v1x, v1y)\n    val mag2 = hypot(v2x, v2y)\n    if (mag1 * mag2 == 0.0) return 180.0\n    val cos = (dot / (mag1 * mag2)).coerceIn(-1.0, 1.0)\n    return Math.toDegrees(acos(cos))\n}\n```\n\nFor a squat, `B` is the knee. `A` is the hip. `C` is the ankle. Same helper, different joints, for every other lift:\n\n| Lift | Primary angle | What “down” looks like | \n|---|---|---|\n| Squat | Knee + hip | Knee `< 105°` and hip`< 115°` | \n| Push-up | Elbow + body line | Elbow `< 95°` ; hips must stay a plank | \n| Deadlift | Hip hinge + knee | Hip `< 115°` into the pull; lockout`> 165°` | \n| Pull-up | Elbow + nose vs wrist | Chin-over-bar or elbow `< 75°` , then a dead hang | \n| Curl | Elbow + upper-arm swing | Peak `< 65°` ; swing`> 35°` is cheating | \n| Plank | Shoulder–hip–ankle | Hold `155–195°` or we flag sag / pike | \n| Running | Alternating ankles | Cadence in steps/min; cue if it drops under 155 SPM | \n\nThe overlay draws the skeleton and the live angles on screen. You can ignore the HUD. The coach is in your ear.\n\nA vision model that shouts “rep!” on a single frame will double-count, miss lockouts, and award half-reps.\n\nEach exercise is a tiny automaton. For squats:\n\n```\nready ──(knee < 105° AND hip < 115°)──► down\ndown  ──(knee > 155° AND hip > 155°)──► up   (+1 rep)\n```\n\nA rep counts **only** on the rising edge - you went down, then you stood back up. The same angles drive form checks:\n\nSpoken cues have a **2.5 second cooldown**. The coach is allowed to be useful. It is not allowed to nag.\n\nPlank is the exception: there is no rep. We score the hold as `good-seconds / total-seconds`. Running scores cadence, not reps.\n\nThat whole layer is the **Reflex Engine**. It is always on. It does not wait for Gemini.\n\nGyms are hostile to pose models. Someone walks behind you, MediaPipe finds a second body, and suddenly you have “done” three of *their* squats.\n\nWe run the landmarker with multiple poses, then lock onto **you**.\n\nThe HUD says `Gym Shield: Passerby Filtered • Holding Rep State`. That line exists because we lost sets to strangers before we wrote it.\n\nVideo never leaves the phone. Landmarks never leave the phone. History is a Room table on device:\n\n```\n@Entity(tableName = \"workout_history\")\ndata class WorkoutSession(\n    val date: Date,\n    val exerciseType: String,\n    val exerciseName: String,\n    val reps: Int,\n    val score: Int,          // form accuracy 0–100\n    val calories: Int = 0,\n    val avgHeartRate: Int = 0,\n    val weightKg: Double = 0.0\n)\n```\n\nWeekly goals and daily challenges read that table with a date-range query. The dashboard calendar, streak, and “200 reps this week” bar are all local.\n\nThe only time we talk to a cloud model is **after** you finish a set, and only if you pasted a Gemini API key. The payload is a paragraph of numbers:\n\n```\nExercise: Squats\nCompleted Reps: 12\nForm Accuracy Score: 88%\nWeight Lifted: 60 kg (Total Volume: 720 kg)\nPrimary Biomechanical Fault: Keep your chest up!\nKey Measured Joint Angle: 82°\n```\n\nNo frame. No landmark list. No face. About 4 KB of text.\n\nIf there is no key, or the request fails, a rule-based report still appears: diagnosis, three cues, two mobility drills. The set is never held hostage by a network.\n\nMid-set, a language model is the wrong tool. After the set, it is the right one.\n\n`GeminiService` calls **Gemini 2.5 Flash**, with **1.5 Flash** as a fallback. We ask for strict JSON, then strip markdown fences if the model gets festive.\n\n**Biomechanics report**\n\n```\n{\n  \"diagnosis\": \"…root cause in two sentences…\",\n  \"correctiveCues\": [\"cue 1\", \"cue 2\", \"cue 3\"],\n  \"accessoryDrills\": [\"drill 1\", \"drill 2\"]\n}\n```\n\n**Recovery protocol**\n\n```\n{\n  \"recoveryWindowHours\": 36,\n  \"strainAssessment\": \"…\",\n  \"proteinGrams\": 35,\n  \"hydrationMl\": 800,\n  \"activeStretches\": [\"…\", \"…\", \"…\"]\n}\n```\n\nThe workout summary bottom sheet shows both. Tap **Hear Coach Review** and ElevenLabs reads the diagnosis plus the cues in the persona you picked.\n\nThat is the multimodal loop people actually feel: *vision counted the set → numbers went to Gemini → voice spoke the plan.*\n\nWhat we did **not** do: send video to Gemini. The model is a sports scientist looking at a stat sheet, not a camera operator.\n\nAndroid TTS can say “rep eight.” It cannot make you believe a coach is in the room.\n\nWe use **ElevenLabs `eleven_turbo_v2_5`** with three personas:\n\n| Persona | Job | When you want them | \n|---|---|---|\n| **Coach Marcus** | Drill sergeant | Last two reps of a hard set | \n| **Coach Maya** | Mindful guide | Breath, tempo, form | \n| **Coach Alex** | Olympian | Bar path and lockout language | \n\nThe latency lesson was simple: if you synthesize *“Great squat depth!”* on every rep, you will miss the next one.\n\nSo the voice stack is cache-first, not stream-first.\n\n`voiceId + phrase` (MD5).`cacheDir/elevenlabs_cache/<hash>.mp3` exists, play it. That is the 0 ms path for counts and stock cues.`TextToSpeech`\nWe download a finished clip rather than streaming chunks. First-time phrases wait on the network (8 s connect / 12 s read timeouts). Repeated gym language - *“Good push-up,” “Rest complete,” “Keep your chest up”* - is instant after the first hit.\n\nOffline, the coach gets flatter. It does not go silent.\n\n`VoiceCommandManager` wraps Android `SpeechRecognizer` and keeps listening in a restart loop.\n\n| You say | FitGuru does | \n|---|---|\n| “finish set” / “end workout” | Stops the camera loop, opens the Gemini summary | \n| “smart rest” / “take a break” | Opens a rest timer that reads current heart rate | \n| “form check” / “coach help” | Speaks a form cue for the current lift | \n| “switch to squat” | Changes the state machine without leaving the HUD | \n| “weight 40” | Sets load for volume math | \n\nTwo guardrails we had to add after the first noisy gym test:\n\nVision sees joints. It cannot see that you are still at 168 BPM.\n\nA Wear OS / Galaxy Watch client can push BPM on the Data Layer path `/heart_rate`. `DataLayerListenerService` broadcasts that into the live workout HUD.\n\nThe **Smart Rest** sheet uses it in a deliberately simple way: if heart rate is already at or under 120, we say you are ready; if not, we tell you to breathe and we let you add 30 seconds. When the timer hits zero, the coach says *“Rest complete.”*\n\nHeart rate also rides along on the `WorkoutSession` into Gemini’s recovery prompt. The watch is optional. The camera coach does not depend on it.\n\nA first-time lifter does not only need a squat counter. They need to know which pin on the lat pulldown is theirs.\n\n`gym_catalog.json` is a 20-machine guide: lat pulldown, seated row, chest press, shoulder press, leg press / extension / curl, cable, Smith, squat rack, dumbbells, bench, pec deck, treadmill, bike, rower, mat, deadlift platform, pull-up bar, running cadence.\n\nEach card has *what it is*, muscle labels, setup steps, execution, common mistakes, and a camera tip. If that movement has a state machine, the card launches live coaching.\n\nOn the home screen:\n\nProfile and history are the same database, not a cloud account.\n\nJudges should not need Android Studio.\n\nThe **web companion** is a Node/Express app with MediaPipe in the page: webcam, skeleton overlay, the same cue dictionary, a wearable BPM simulator. It deploys as a Docker container (DigitalOcean App Platform spec is in the repo).\n\niOS is a SwiftUI client: Apple Vision (`VNDetectHumanBodyPoseRequest`) on device, the same ElevenLabs + Gemini services, gym catalog bundled, history in local storage.\n\nAndroid is the deepest implementation. The idea is the same everywhere: **pose on device, language after the set, voice as the interface.**\n\nWe wanted Gemma on-device for mid-set language. Running a ~1.3 GB model while CameraX + MediaPipe were live caused **native memory crashes**.\n\n`PoseLandmarkerHelper` now has `pause()` / `resume()` so we can tear down the landmarker, free the buffers, run a heavy inference, and come back. In the current Android build, mid-set “Gemma AI” lines are a **templated cognitive layer** (fault type → a longer cue) with a short thinking indicator - not a 2B model sitting in RAM next to the camera. The *real* LLM call is Gemini, and it happens **after** the set, on a JSON summary.\n\nThe fallback is the Reflex Engine. If the clever layer is thinking, or dead, reps still count.\n\nFirst pass at watch data used the Sensor SDK. Builds failed and the APIs were inconsistent. We switched to the **Wearable Data Layer** (`/heart_rate` → local broadcast). Heart rate shows in the HUD and Smart Rest when a watch is paired. It is optional, not a hard dependency.\n\n`knee < 105°` is a good squat for a lot of people and a terrible one for others. Femur length, camera height, and stance change every angle. We tightened the squat to **both** knee and hip, required a full stand-up before counting, and added the 2.5 s cue cooldown so a noisy angle does not become a TED talk. Per-user calibration is still the honest next step. We have not shipped it.\n\nFirst ElevenLabs + SpeechRecognizer pairing: the mic heard “finish set” from the speaker. Debounce + “do not listen while speaking” fixed it.\n\nBefore Gym Shield, a walker behind you could own the landmarker. Prominence scoring + a 45-frame occlusion hold was the difference between a demo and a gym app.\n\nWe started from **Google’s official MediaPipe Pose Landmarker Android sample** and kept the CameraX + landmarker wiring. That saved the week we would have spent on YUV buffers.\n\nEverything above that sample is ours: exercise state machines, Gym Shield, Room history, challenges, gym catalog, Gemini debrief, ElevenLabs cache and personas, voice commands, Smart Rest, analytics, web companion, and the iOS client.\n\nCredit the sample. Do not confuse it with the product.\n\n| **Resource** | **Link / Command** | \n|---|---|\n| **Source Code** | [https://github.com/atharvachaudhari1/innovhacks4.0_arka](https://github.com/atharvachaudhari1/innovhacks4.0_arka) | \n| **Android APK** | `adb install apk/FitGuru-Innovhacks4.apk` (API 29+) | \n| **Web Companion** | `cd web-companion && npm install && npm start` →`http://localhost:8080` | \n| **Live Demo** | [https://web-landing-76ktyl4c8-bharat7.vercel.app/](https://web-landing-76ktyl4c8-bharat7.vercel.app/) | \n| **Technical Article:** | [Read On Devfolio](https://devfolio.co/projects/fitguru-297b) | \n\nOptional keys (Settings / SharedPreferences): Gemini for the post-set report, ElevenLabs for studio voices. Neither is required to count a squat.\n\n**Stack, in one line:** Kotlin · CameraX · MediaPipe Pose · joint-angle Reflex Engine · Room · Android SpeechRecognizer · ElevenLabs turbo + TTS fallback · Gemini 2.5 Flash · Wear Data Layer · Node web companion · SwiftUI + Vision.\n\nBuilt by **Team ARKA** - **Allan Fernandes** and **Atharva Chaudhari** - for **Innovhacks 4.0**.\n\nThe coach lives in your pocket. The video stays there too.", "url": "https://wpnews.pro/news/fitguru-building-a-private-on-device-ai-workout-coach-in-kotlin-multimodel-ai", "canonical_source": "https://dev.to/atharvachaudhari1/fitguru-building-a-private-on-device-ai-workout-coach-in-kotlinmultimodel-ai-2ngg", "published_at": "2026-10-11 04:37:23+00:00", "updated_at": "2026-10-11 04:49:54.068634+00:00", "lang": "en", "topics": ["artificial-intelligence", "generative-ai", "computer-vision", "ai-agents", "ai-products"], "entities": ["FitGuru", "Team ARKA", "Innovhacks 4.0", "MediaPipe", "Gemini", "ElevenLabs", "Major League Hacking", "Kotlin"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/fitguru-building-a-private-on-device-ai-workout-coach-in-kotlin-multimodel-ai", "markdown": "https://wpnews.pro/news/fitguru-building-a-private-on-device-ai-workout-coach-in-kotlin-multimodel-ai.md", "text": "https://wpnews.pro/news/fitguru-building-a-private-on-device-ai-workout-coach-in-kotlin-multimodel-ai.txt", "jsonld": "https://wpnews.pro/news/fitguru-building-a-private-on-device-ai-workout-coach-in-kotlin-multimodel-ai.jsonld"}}