published: false
description: "We built a phone-side workout coach at Innovhacks 4.0 - MediaPipe pose, joint-angle state machines, Gemini debriefs, and ElevenLabs voice - without sending your camera feed to the cloud."
tags: ai, genai, agentaichallenge, kotlin
Innovhacks 4.0 · Team ARKA · Android (Kotlin) · Multimodal AI
You are three reps into a hard set. You cannot look at a screen. You also cannot tell if your squat is deep enough, if your chest collapsed, or if that last lockout even counted.
Most “smart” fitness apps solve this by shipping your camera to a server. That is a lot to ask for video of your body in a gym.
At Innovhacks 4.0, Team ARKA built FitGuru to test a narrower idea:
Can a phone watch your form, count the set, talk you through mistakes, and debrief you afterwards - without the video ever leaving the device?
The short answer: yes, if you split the coach into two brains. A cheap Reflex Engine that runs on every camera frame. And a Cognitive Engine that only wakes up when the set is over, and only sees numbers.
Demo:https://drive.google.com/file/d/1yyQeH4TjNK_Y5u2S9msBnvTxCkJeANAF/view?usp=sharing
Code: github.com/atharvachaudhari1/innovhacks4.0_arka
APK: apk/FitGuru-Innovhacks4.apk
FitGuru was our submission to Innovhacks 4.0 through Major League Hacking (MLH). We had a weekend and one hard rule (the video never leaves the phone), which is why the architecture ended up as two brains instead of one giant model.
We also used two of the tools MLH hackathons often put in front of builders:
If the camera feed leaves the phone, we already lost.
That one rule decided the architecture:
| Must stay on the phone | Allowed off the phone (opt-in) |
|---|---|
| Camera frames | Nothing visual |
| 33 pose landmarks | Nothing visual |
| Rep counting + form cues | Structured JSON of angles, reps, score, heart rate |
| Workout history (Room) | Gemini debrief text |
| Offline TTS fallback | ElevenLabs audio, if the user configured a key |
No account. No cloud video pipeline. If the network dies mid-session, the coach still counts reps and still talks.
Open the app, pick a lift, put the phone where it can see you, and start the set.
Live form analysis for 18 exercises
Squats, push-ups, deadlifts, pull-ups, chin-ups, lunges, plank, bicep curls, running cadence, bench / chest press, shoulder press, leg press / extension / curl, lat pulldown, seated row, triceps pushdown.
Automatic rep counting using joint-angle state machines. A squat is not a guess. A rep exists only when the knee and hip cross a “down” threshold and then return past an “up” threshold.
Spoken corrections so you never look at the screen mid-set. Short cues during the set. A longer spoken review after you say “finish set.”
Hands-free voice commands - finish set, smart rest, form check, switch to deadlift, weight 40.
Local history - reps, form score, volume, calories, average heart rate - stored in Room. Daily challenges and a weekly rep goal live on the dashboard.
A 20-machine gym guide for beginners: what the machine does, how to set it up, common mistakes, and whether FitGuru can coach that movement live.
A web companion with the same MediaPipe loop in the browser, plus an iOS SwiftUI client that uses Apple Vision.
The tempting hackathon move is: stream frames to an LLM and ask it “is this squat good?”
That fails in three ways at once.
So FitGuru is multimodal by role, not by dumping every sensor into one prompt.
Camera ──► MediaPipe Pose (on-device)
│
▼
33 landmarks / frame
│
▼
Reflex Engine (always on)
• joint angles
• state machine (ready → down → up)
• form faults + 2.5s cue cooldown
• Gym Shield (lock onto you, ignore walkers)
│
┌──────────┴──────────┐
▼ ▼
ElevenLabs / TTS Room (history)
│
▼ only after the set
Gemini 2.5 Flash
(JSON of numbers, never video)
│
▼
Spoken debrief + recovery protocol
Reflex is trigonometry and thresholds. It is cheap enough to run on every frame.
Cognitive is Gemini. It sees a session summary, not a pixel.
That split is the product.
The Android camera path is CameraX → MediaPipe Pose Landmarker (pose_landmarker_full.task) in LIVE_STREAM mode. Each result is 33 normalized 3D landmarks: nose, shoulders, elbows, wrists, hips, knees, ankles, toes.
We never ask the model “was that a squat.” We compute the interior angle at a joint.
// angle at landmark B, formed by A–B–C
fun calculateAngle(a: Landmark, b: Landmark, c: Landmark): Double {
val v1x = a.x() - b.x(); val v1y = a.y() - b.y()
val v2x = c.x() - b.x(); val v2y = c.y() - b.y()
val dot = v1x * v2x + v1y * v2y
val mag1 = hypot(v1x, v1y)
val mag2 = hypot(v2x, v2y)
if (mag1 * mag2 == 0.0) return 180.0
val cos = (dot / (mag1 * mag2)).coerceIn(-1.0, 1.0)
return Math.toDegrees(acos(cos))
}
For a squat, B is the knee. A is the hip. C is the ankle. Same helper, different joints, for every other lift:
| Lift | Primary angle | What “down” looks like |
|---|---|---|
| Squat | Knee + hip | Knee < 105° and hip< 115° |
| Push-up | Elbow + body line | Elbow < 95° ; hips must stay a plank |
| Deadlift | Hip hinge + knee | Hip < 115° into the pull; lockout> 165° |
| Pull-up | Elbow + nose vs wrist | Chin-over-bar or elbow < 75° , then a dead hang |
| Curl | Elbow + upper-arm swing | Peak < 65° ; swing> 35° is cheating |
| Plank | Shoulder–hip–ankle | Hold 155–195° or we flag sag / pike |
| Running | Alternating ankles | Cadence in steps/min; cue if it drops under 155 SPM |
The overlay draws the skeleton and the live angles on screen. You can ignore the HUD. The coach is in your ear.
A vision model that shouts “rep!” on a single frame will double-count, miss lockouts, and award half-reps.
Each exercise is a tiny automaton. For squats:
ready ──(knee < 105° AND hip < 115°)──► down
down ──(knee > 155° AND hip > 155°)──► up (+1 rep)
A rep counts only on the rising edge - you went down, then you stood back up. The same angles drive form checks:
Spoken cues have a 2.5 second cooldown. The coach is allowed to be useful. It is not allowed to nag.
Plank is the exception: there is no rep. We score the hold as good-seconds / total-seconds. Running scores cadence, not reps.
That whole layer is the Reflex Engine. It is always on. It does not wait for Gemini.
Gyms are hostile to pose models. Someone walks behind you, MediaPipe finds a second body, and suddenly you have “done” three of their squats.
We run the landmarker with multiple poses, then lock onto you.
The HUD says Gym Shield: Passerby Filtered • Holding Rep State. That line exists because we lost sets to strangers before we wrote it.
Video never leaves the phone. Landmarks never leave the phone. History is a Room table on device:
@Entity(tableName = "workout_history")
data class WorkoutSession(
val date: Date,
val exerciseType: String,
val exerciseName: String,
val reps: Int,
val score: Int, // form accuracy 0–100
val calories: Int = 0,
val avgHeartRate: Int = 0,
val weightKg: Double = 0.0
)
Weekly goals and daily challenges read that table with a date-range query. The dashboard calendar, streak, and “200 reps this week” bar are all local.
The only time we talk to a cloud model is after you finish a set, and only if you pasted a Gemini API key. The payload is a paragraph of numbers:
Exercise: Squats
Completed Reps: 12
Form Accuracy Score: 88%
Weight Lifted: 60 kg (Total Volume: 720 kg)
Primary Biomechanical Fault: Keep your chest up!
Key Measured Joint Angle: 82°
No frame. No landmark list. No face. About 4 KB of text.
If there is no key, or the request fails, a rule-based report still appears: diagnosis, three cues, two mobility drills. The set is never held hostage by a network.
Mid-set, a language model is the wrong tool. After the set, it is the right one.
GeminiService calls Gemini 2.5 Flash, with 1.5 Flash as a fallback. We ask for strict JSON, then strip markdown fences if the model gets festive.
Biomechanics report
{
"diagnosis": "…root cause in two sentences…",
"correctiveCues": ["cue 1", "cue 2", "cue 3"],
"accessoryDrills": ["drill 1", "drill 2"]
}
Recovery protocol
{
"recoveryWindowHours": 36,
"strainAssessment": "…",
"proteinGrams": 35,
"hydrationMl": 800,
"activeStretches": ["…", "…", "…"]
}
The workout summary bottom sheet shows both. Tap Hear Coach Review and ElevenLabs reads the diagnosis plus the cues in the persona you picked.
That is the multimodal loop people actually feel: vision counted the set → numbers went to Gemini → voice spoke the plan.
What we did not do: send video to Gemini. The model is a sports scientist looking at a stat sheet, not a camera operator.
Android TTS can say “rep eight.” It cannot make you believe a coach is in the room.
We use ElevenLabs eleven_turbo_v2_5 with three personas:
| Persona | Job | When you want them |
|---|---|---|
| Coach Marcus | Drill sergeant | Last two reps of a hard set |
| Coach Maya | Mindful guide | Breath, tempo, form |
| Coach Alex | Olympian | Bar path and lockout language |
The latency lesson was simple: if you synthesize “Great squat depth!” on every rep, you will miss the next one.
So the voice stack is cache-first, not stream-first.
voiceId + phrase (MD5).cacheDir/elevenlabs_cache/<hash>.mp3 exists, play it. That is the 0 ms path for counts and stock cues.TextToSpeech
We download a finished clip rather than streaming chunks. First-time phrases wait on the network (8 s connect / 12 s read timeouts). Repeated gym language - “Good push-up,” “Rest complete,” “Keep your chest up” - is instant after the first hit.
Offline, the coach gets flatter. It does not go silent.
VoiceCommandManager wraps Android SpeechRecognizer and keeps listening in a restart loop.
| You say | FitGuru does |
|---|---|
| “finish set” / “end workout” | Stops the camera loop, opens the Gemini summary |
| “smart rest” / “take a break” | Opens a rest timer that reads current heart rate |
| “form check” / “coach help” | Speaks a form cue for the current lift |
| “switch to squat” | Changes the state machine without leaving the HUD |
| “weight 40” | Sets load for volume math |
Two guardrails we had to add after the first noisy gym test:
Vision sees joints. It cannot see that you are still at 168 BPM.
A Wear OS / Galaxy Watch client can push BPM on the Data Layer path /heart_rate. DataLayerListenerService broadcasts that into the live workout HUD.
The Smart Rest sheet uses it in a deliberately simple way: if heart rate is already at or under 120, we say you are ready; if not, we tell you to breathe and we let you add 30 seconds. When the timer hits zero, the coach says “Rest complete.”
Heart rate also rides along on the WorkoutSession into Gemini’s recovery prompt. The watch is optional. The camera coach does not depend on it.
A first-time lifter does not only need a squat counter. They need to know which pin on the lat pulldown is theirs.
gym_catalog.json is a 20-machine guide: lat pulldown, seated row, chest press, shoulder press, leg press / extension / curl, cable, Smith, squat rack, dumbbells, bench, pec deck, treadmill, bike, rower, mat, deadlift platform, pull-up bar, running cadence.
Each card has what it is, muscle labels, setup steps, execution, common mistakes, and a camera tip. If that movement has a state machine, the card launches live coaching.
On the home screen:
Profile and history are the same database, not a cloud account.
Judges should not need Android Studio.
The web companion is a Node/Express app with MediaPipe in the page: webcam, skeleton overlay, the same cue dictionary, a wearable BPM simulator. It deploys as a Docker container (DigitalOcean App Platform spec is in the repo).
iOS is a SwiftUI client: Apple Vision (VNDetectHumanBodyPoseRequest) on device, the same ElevenLabs + Gemini services, gym catalog bundled, history in local storage.
Android is the deepest implementation. The idea is the same everywhere: pose on device, language after the set, voice as the interface.
We wanted Gemma on-device for mid-set language. Running a ~1.3 GB model while CameraX + MediaPipe were live caused native memory crashes.
PoseLandmarkerHelper now has () / resume() so we can tear down the landmarker, free the buffers, run a heavy inference, and come back. In the current Android build, mid-set “Gemma AI” lines are a templated cognitive layer (fault type → a longer cue) with a short thinking indicator - not a 2B model sitting in RAM next to the camera. The real LLM call is Gemini, and it happens after the set, on a JSON summary.
The fallback is the Reflex Engine. If the clever layer is thinking, or dead, reps still count.
First pass at watch data used the Sensor SDK. Builds failed and the APIs were inconsistent. We switched to the Wearable Data Layer (/heart_rate → local broadcast). Heart rate shows in the HUD and Smart Rest when a watch is paired. It is optional, not a hard dependency.
knee < 105° is a good squat for a lot of people and a terrible one for others. Femur length, camera height, and stance change every angle. We tightened the squat to both knee and hip, required a full stand-up before counting, and added the 2.5 s cue cooldown so a noisy angle does not become a TED talk. Per-user calibration is still the honest next step. We have not shipped it.
First ElevenLabs + SpeechRecognizer pairing: the mic heard “finish set” from the speaker. Debounce + “do not listen while speaking” fixed it.
Before Gym Shield, a walker behind you could own the landmarker. Prominence scoring + a 45-frame occlusion hold was the difference between a demo and a gym app.
We started from Google’s official MediaPipe Pose Landmarker Android sample and kept the CameraX + landmarker wiring. That saved the week we would have spent on YUV buffers.
Everything above that sample is ours: exercise state machines, Gym Shield, Room history, challenges, gym catalog, Gemini debrief, ElevenLabs cache and personas, voice commands, Smart Rest, analytics, web companion, and the iOS client.
Credit the sample. Do not confuse it with the product.
| Resource | Link / Command |
|---|---|
| Source Code | https://github.com/atharvachaudhari1/innovhacks4.0_arka |
| Android APK | adb install apk/FitGuru-Innovhacks4.apk (API 29+) |
| Web Companion | cd web-companion && npm install && npm start →http://localhost:8080 |
| Live Demo | https://web-landing-76ktyl4c8-bharat7.vercel.app/ |
| Technical Article: | Read On Devfolio |
Optional keys (Settings / SharedPreferences): Gemini for the post-set report, ElevenLabs for studio voices. Neither is required to count a squat.
Stack, in one line: Kotlin · CameraX · MediaPipe Pose · joint-angle Reflex Engine · Room · Android SpeechRecognizer · ElevenLabs turbo + TTS fallback · Gemini 2.5 Flash · Wear Data Layer · Node web companion · SwiftUI + Vision.
Built by Team ARKA - Allan Fernandes and Atharva Chaudhari - for Innovhacks 4.0.
The coach lives in your pocket. The video stays there too.