FitGuru: Building a Private, On-Device AI Workout Coach in Kotlin(Multimodel AI) Team ARKA built FitGuru, an Android workout coach that runs MediaPipe pose estimation entirely on-device, using joint-angle state machines for rep counting and form cues while only sending structured JSON (angles, reps, score, heart rate) off the phone for optional Gemini debriefs and ElevenLabs voice. The architecture splits into an always-on Reflex Engine for per-frame analysis and a Cognitive Engine that activates after a set, keeping camera frames and 33 pose landmarks local. The project was submitted at Innovhacks 4.0 and supports live form analysis for 18 exercises. published: false description: "We built a phone-side workout coach at Innovhacks 4.0 - MediaPipe pose, joint-angle state machines, Gemini debriefs, and ElevenLabs voice - without sending your camera feed to the cloud." tags: ai, genai, agentaichallenge, kotlin Innovhacks 4.0 · Team ARKA · Android Kotlin · Multimodal AI You are three reps into a hard set. You cannot look at a screen. You also cannot tell if your squat is deep enough, if your chest collapsed, or if that last lockout even counted. Most “smart” fitness apps solve this by shipping your camera to a server. That is a lot to ask for video of your body in a gym. At Innovhacks 4.0 , Team ARKA built FitGuru to test a narrower idea: Can a phone watch your form, count the set, talk you through mistakes, and debrief you afterwards - without the video ever leaving the device? The short answer: yes, if you split the coach into two brains. A cheap Reflex Engine that runs on every camera frame. And a Cognitive Engine that only wakes up when the set is over, and only sees numbers. Demo: https://drive.google.com/file/d/1yyQeH4TjNK Y5u2S9msBnvTxCkJeANAF/view?usp=sharing https://drive.google.com/file/d/1yyQeH4TjNK Y5u2S9msBnvTxCkJeANAF/view?usp=sharing Code: github.com/atharvachaudhari1/innovhacks4.0 arka https://github.com/atharvachaudhari1/innovhacks4.0 arka APK: apk/FitGuru-Innovhacks4.apk FitGuru was our submission to Innovhacks 4.0 through Major League Hacking MLH . We had a weekend and one hard rule the video never leaves the phone , which is why the architecture ended up as two brains instead of one giant model. We also used two of the tools MLH hackathons often put in front of builders: If the camera feed leaves the phone, we already lost. That one rule decided the architecture: | Must stay on the phone | Allowed off the phone opt-in | |---|---| | Camera frames | Nothing visual | | 33 pose landmarks | Nothing visual | | Rep counting + form cues | Structured JSON of angles, reps, score, heart rate | | Workout history Room | Gemini debrief text | | Offline TTS fallback | ElevenLabs audio, if the user configured a key | No account. No cloud video pipeline. If the network dies mid-session, the coach still counts reps and still talks. Open the app, pick a lift, put the phone where it can see you, and start the set. Live form analysis for 18 exercises Squats, push-ups, deadlifts, pull-ups, chin-ups, lunges, plank, bicep curls, running cadence, bench / chest press, shoulder press, leg press / extension / curl, lat pulldown, seated row, triceps pushdown. Automatic rep counting using joint-angle state machines. A squat is not a guess. A rep exists only when the knee and hip cross a “down” threshold and then return past an “up” threshold. Spoken corrections so you never look at the screen mid-set. Short cues during the set. A longer spoken review after you say “finish set.” Hands-free voice commands - finish set , smart rest , form check , switch to deadlift , weight 40 . Local history - reps, form score, volume, calories, average heart rate - stored in Room. Daily challenges and a weekly rep goal live on the dashboard. A 20-machine gym guide for beginners: what the machine does, how to set it up, common mistakes, and whether FitGuru can coach that movement live. A web companion with the same MediaPipe loop in the browser, plus an iOS SwiftUI client that uses Apple Vision. The tempting hackathon move is: stream frames to an LLM and ask it “is this squat good?” That fails in three ways at once. So FitGuru is multimodal by role , not by dumping every sensor into one prompt. Camera ──► MediaPipe Pose on-device │ ▼ 33 landmarks / frame │ ▼ Reflex Engine always on • joint angles • state machine ready → down → up • form faults + 2.5s cue cooldown • Gym Shield lock onto you, ignore walkers │ ┌──────────┴──────────┐ ▼ ▼ ElevenLabs / TTS Room history │ ▼ only after the set Gemini 2.5 Flash JSON of numbers, never video │ ▼ Spoken debrief + recovery protocol Reflex is trigonometry and thresholds. It is cheap enough to run on every frame. Cognitive is Gemini. It sees a session summary, not a pixel. That split is the product. The Android camera path is CameraX → MediaPipe Pose Landmarker pose landmarker full.task in LIVE STREAM mode. Each result is 33 normalized 3D landmarks: nose, shoulders, elbows, wrists, hips, knees, ankles, toes. We never ask the model “was that a squat.” We compute the interior angle at a joint. // angle at landmark B, formed by A–B–C fun calculateAngle a: Landmark, b: Landmark, c: Landmark : Double { val v1x = a.x - b.x ; val v1y = a.y - b.y val v2x = c.x - b.x ; val v2y = c.y - b.y val dot = v1x v2x + v1y v2y val mag1 = hypot v1x, v1y val mag2 = hypot v2x, v2y if mag1 mag2 == 0.0 return 180.0 val cos = dot / mag1 mag2 .coerceIn -1.0, 1.0 return Math.toDegrees acos cos } For a squat, B is the knee. A is the hip. C is the ankle. Same helper, different joints, for every other lift: | Lift | Primary angle | What “down” looks like | |---|---|---| | Squat | Knee + hip | Knee < 105° and hip < 115° | | Push-up | Elbow + body line | Elbow < 95° ; hips must stay a plank | | Deadlift | Hip hinge + knee | Hip < 115° into the pull; lockout 165° | | Pull-up | Elbow + nose vs wrist | Chin-over-bar or elbow < 75° , then a dead hang | | Curl | Elbow + upper-arm swing | Peak < 65° ; swing 35° is cheating | | Plank | Shoulder–hip–ankle | Hold 155–195° or we flag sag / pike | | Running | Alternating ankles | Cadence in steps/min; cue if it drops under 155 SPM | The overlay draws the skeleton and the live angles on screen. You can ignore the HUD. The coach is in your ear. A vision model that shouts “rep ” on a single frame will double-count, miss lockouts, and award half-reps. Each exercise is a tiny automaton. For squats: ready ── knee < 105° AND hip < 115° ──► down down ── knee 155° AND hip 155° ──► up +1 rep A rep counts only on the rising edge - you went down, then you stood back up. The same angles drive form checks: Spoken cues have a 2.5 second cooldown . The coach is allowed to be useful. It is not allowed to nag. Plank is the exception: there is no rep. We score the hold as good-seconds / total-seconds . Running scores cadence, not reps. That whole layer is the Reflex Engine . It is always on. It does not wait for Gemini. Gyms are hostile to pose models. Someone walks behind you, MediaPipe finds a second body, and suddenly you have “done” three of their squats. We run the landmarker with multiple poses, then lock onto you . The HUD says Gym Shield: Passerby Filtered • Holding Rep State . That line exists because we lost sets to strangers before we wrote it. Video never leaves the phone. Landmarks never leave the phone. History is a Room table on device: @Entity tableName = "workout history" data class WorkoutSession val date: Date, val exerciseType: String, val exerciseName: String, val reps: Int, val score: Int, // form accuracy 0–100 val calories: Int = 0, val avgHeartRate: Int = 0, val weightKg: Double = 0.0 Weekly goals and daily challenges read that table with a date-range query. The dashboard calendar, streak, and “200 reps this week” bar are all local. The only time we talk to a cloud model is after you finish a set, and only if you pasted a Gemini API key. The payload is a paragraph of numbers: Exercise: Squats Completed Reps: 12 Form Accuracy Score: 88% Weight Lifted: 60 kg Total Volume: 720 kg Primary Biomechanical Fault: Keep your chest up Key Measured Joint Angle: 82° No frame. No landmark list. No face. About 4 KB of text. If there is no key, or the request fails, a rule-based report still appears: diagnosis, three cues, two mobility drills. The set is never held hostage by a network. Mid-set, a language model is the wrong tool. After the set, it is the right one. GeminiService calls Gemini 2.5 Flash , with 1.5 Flash as a fallback. We ask for strict JSON, then strip markdown fences if the model gets festive. Biomechanics report { "diagnosis": "…root cause in two sentences…", "correctiveCues": "cue 1", "cue 2", "cue 3" , "accessoryDrills": "drill 1", "drill 2" } Recovery protocol { "recoveryWindowHours": 36, "strainAssessment": "…", "proteinGrams": 35, "hydrationMl": 800, "activeStretches": "…", "…", "…" } The workout summary bottom sheet shows both. Tap Hear Coach Review and ElevenLabs reads the diagnosis plus the cues in the persona you picked. That is the multimodal loop people actually feel: vision counted the set → numbers went to Gemini → voice spoke the plan. What we did not do: send video to Gemini. The model is a sports scientist looking at a stat sheet, not a camera operator. Android TTS can say “rep eight.” It cannot make you believe a coach is in the room. We use ElevenLabs eleven turbo v2 5 with three personas: | Persona | Job | When you want them | |---|---|---| | Coach Marcus | Drill sergeant | Last two reps of a hard set | | Coach Maya | Mindful guide | Breath, tempo, form | | Coach Alex | Olympian | Bar path and lockout language | The latency lesson was simple: if you synthesize “Great squat depth ” on every rep, you will miss the next one. So the voice stack is cache-first, not stream-first. voiceId + phrase MD5 . cacheDir/elevenlabs cache/