{"slug": "how-our-top-11-hophacks-project-turned-noisy-mediapipe-pose-data-into-reliable", "title": "How our Top 11 HopHacks project turned noisy MediaPipe pose data into reliable rehabilitation metrics for adaptive AI planning", "summary": "A developer on the Mendly team built a computer-vision pipeline for stroke rehabilitation that converts MediaPipe Pose landmarks into joint angles, range-of-motion measurements, and valid-repetition counts for an adaptive AI planning layer. The project, a Top 11 finalist among 87 entries at HopHacks and winner of Best Use of DigitalOcean, treats failed pose inference as an invalid observation rather than substituting zero or the last known angle, preventing false movement data from propagating into performance metrics. The team also addressed normalized-coordinate and tracking-loss issues that made raw camera observations unreliable.", "body_md": "What if a rehabilitation exercise could understand **how** a patient moved, instead of only asking whether they finished?\n\nThat was one of the questions behind **Mendly**, an AI-powered stroke rehabilitation platform my team built during HopHacks.\n\nBy the end of the hackathon, Mendly became a **Top 11 finalist out of 87 projects** and won **Best Use of DigitalOcean**.\n\nI worked mainly on the computer vision side of the project: pose tracking, joint-angle and range-of-motion measurement, valid-repetition detection, tracking-loss handling, camera reliability, and the layer that turns movement into structured performance data for our adaptive AI pipeline.\n\nThe part I found most interesting was not simply getting MediaPipe to recognize a person.\n\nIt was figuring out how to turn imperfect camera observations into data that the rest of the system could actually trust.\n\nThe pipeline looked roughly like this:\n\n```\nCamera\n  ↓\nPose Detection\n  ↓\nJoint Angles / Range of Motion\n  ↓\nReliability Checks\n  ↓\nRepetition Detection\n  ↓\nStructured Performance Data\n  ↓\nFuture Adaptive Planning\n```\n\nThat sounds straightforward on paper.\n\nIt wasn't.\n\nWe used **MediaPipe Pose** to estimate landmarks on the patient's body.\n\nInstead of trying to reason directly about image pixels, we could work with points representing joints such as the shoulder, elbow, wrist, hip, knee, and ankle.\n\nFor example, if we know the positions of the shoulder, elbow, and wrist, we can calculate the angle at the elbow.\n\nGiven two vectors around a joint:\n\n```\nv₁ = Shoulder - Elbow\nv₂ = Wrist - Elbow\n```\n\nwe can calculate the angle between them with the dot product:\n\n```\nθ = arccos((v₁ · v₂) / (|v₁||v₂|))\n```\n\nOnce we have joint angles over time, we can start measuring things like range of motion.\n\nIf an elbow moves from roughly:\n\n```\n155° → 72°\n```\n\nthen the observed range is approximately:\n\n```\n83°\n```\n\nThis is already much more useful than simply knowing that \"the arm moved.\"\n\nBut getting an angle is the easy part.\n\nThe harder question is whether that angle should be trusted.\n\nThis was probably the biggest thing I learned from the project.\n\nAt first, it is easy to think of pose estimation like this:\n\n```\nMediaPipe detected a wrist\n        ↓\ntherefore the wrist position is correct\n```\n\nReal camera input is much messier.\n\nA patient can move partly outside the frame. A joint can become occluded. Lighting can change. Landmarks can jump. The patient can turn sideways. The camera can be positioned badly.\n\nA single bad landmark can propagate through everything after it:\n\n```\nbad landmark\n    ↓\nbad angle\n    ↓\nbad repetition\n    ↓\nbad performance data\n    ↓\nbad future planning context\n```\n\nThat made reliability part of the actual product, not just an implementation detail.\n\nSome of the most interesting bugs appeared when we asked a simple question:\n\n**What should happen when the camera does not know?**\n\nSuppose the tracker sees:\n\n```\nFrame 1: 40°\nFrame 2: 55°\nFrame 3: tracking fails\n```\n\nWhat should Frame 3 become?\n\nThere are two tempting answers.\n\nOne is:\n\n```\n0°\n```\n\nThe other is:\n\n```\n55°\n```\n\nNeither is correct.\n\nIf tracking failed, we do not know where the patient's arm actually was.\n\nTreating the missing value as zero invents a movement that never happened.\n\nKeeping the previous angle is not much better. The patient may have continued moving while the camera lost them.\n\nSo we changed the pipeline to represent failed inference as an **invalid observation**.\n\n```\nPose succeeds\n→ process the measurement\n\nPose fails\n→ mark observation invalid\n→ do not update movement state\n→ do not manufacture an angle\n→ do not count a repetition\n```\n\nThe rule became:\n\n**Unknown is not zero. Unknown stays unknown.**\n\nThat sounds obvious after the fact, but it fixed an important class of false movement data.\n\nAnother issue was caused by normalized coordinates.\n\nMediaPipe gives landmark positions as normalized X and Y values.\n\nBut X is normalized relative to the frame width, while Y is normalized relative to the frame height.\n\nOn a 640×480 camera:\n\n```\nΔx = 0.1 → 64 pixels\nΔy = 0.1 → 48 pixels\n```\n\nThose values are both `0.1` in normalized space, but they do not represent the same physical distance in the image.\n\nIf we calculate geometry as if the normalized X and Y scales are identical, joint angles can become distorted.\n\nWe fixed this by accounting for the actual camera aspect ratio before calculating angles and ray lengths.\n\nAt the same time, we kept the original normalized landmarks for things like checking whether a joint was still inside the frame.\n\nThis was a good example of a bug that was mathematically small but affected everything built on top of it.\n\nAnother thing that looked simple at first was repetition counting.\n\nA naive approach might say:\n\nCount a rep whenever the joint angle crosses a threshold.\n\nThat works until real movement starts looking like this:\n\n```\n155°\n148°\n139°\n121°\n97°\n76°\n81°\n95°\n119°\n142°\n151°\n```\n\nHuman motion does not happen in perfectly clean steps.\n\nThe angle may hover around a threshold, move backward briefly, or contain noise.\n\nSo instead of treating each frame independently, we used a state-based approach.\n\nA complete repetition is closer to:\n\n```\nStart\n  ↓\nMovement begins\n  ↓\nRequired range reached\n  ↓\nHold requirement satisfied, if needed\n  ↓\nMovement reverses\n  ↓\nReturn condition reached\n  ↓\nValid repetition\n```\n\nCrossing one threshold is not automatically a repetition.\n\nA rep has a beginning, progression, target, and completion.\n\nDuring testing, we found another behavior that looked reasonable in code but felt completely wrong from the patient's perspective.\n\nImagine the exercise target is:\n\n```\n5 valid reps\n```\n\nand the patient does:\n\n```\nvalid\nfailed\nfailed\nvalid\nfailed\nfailed\nfailed\n```\n\nThey have completed only:\n\n```\n2 / 5 valid reps\n```\n\nBut an earlier completion rule could still stop the exercise after enough total attempts.\n\nThat meant a patient could fail several repetitions and somehow \"finish\" the exercise.\n\nWe changed that.\n\nAutomatic completion now depends on the number of **valid repetitions**, not the total number of attempts.\n\n```\nvalidRepCount >= targetRepCount\n```\n\nFailed attempts can still be recorded.\n\nThey just do not count toward the target.\n\nSo if the prescription says five valid repetitions, the patient actually needs five valid repetitions.\n\nThis distinction also became important for range-of-motion statistics.\n\nSuppose a patient performs a repetition while the camera is barely tracking the required landmarks.\n\nWe may still want to know that an attempt happened.\n\nBut that does not mean its ROM measurement should be treated as reliable.\n\nEarlier logic could fall back to unreliable repetitions when there were no reliable ones available.\n\nWe changed that behavior.\n\nA low-confidence attempt can remain part of the exercise history, but unreliable motion should not contaminate trusted ROM statistics.\n\nI started thinking about this as two different questions:\n\n```\nDid something happen?\n```\n\nand:\n\n```\nDo we trust this measurement?\n```\n\nThose are not the same thing.\n\nPose tracking also depends heavily on how the patient positions the camera.\n\nWe built a setup checker to decide whether the current view was usable for the exercise.\n\nOne issue was that early bad frames could influence the setup state for too long.\n\nImagine this:\n\n```\nbad setup\nbad setup\nbad setup\nbad setup\n\npatient fixes the camera\n\ngood\ngood\ngood\n```\n\nIf the checker effectively remembers the entire history, the patient can remain stuck even after fixing the problem.\n\nWe changed the logic to focus on a recent rolling window instead.\n\nThat way, old bad setup frames do not permanently punish the patient.\n\nWe also kept checking the environment after the patient reached the Ready state.\n\nOtherwise this could happen:\n\n```\nPatient is visible\n→ Ready\n\nPatient leaves the frame\n→ still Ready\n```\n\nInstead, if the setup becomes unreliable again, Ready can be revoked.\n\nWe also added a timeout.\n\nIf the pose stream is running but the system cannot establish a reliable camera view after around 25 seconds, it stops waiting forever and shows the patient a retry option.\n\nThat turned camera setup from a one-shot gate into something that could actually recover from failure.\n\nAnother bug was less about computer vision and more about product consistency.\n\nAn exercise level could require something like:\n\n```\n12 reps\n70° movement target\n2-second hold\n```\n\nThe tracker knew about the hold.\n\nThe patient did not.\n\nThe UI originally might only say:\n\nRaise your arm out to the side, then lower it slowly.\n\nFrom the patient's point of view, they could perform the movement correctly according to the instruction and still fail the tracker.\n\nWe fixed this by deriving the patient-facing prescription from the same exercise-level configuration used by the tracking logic.\n\nSo the interface could show:\n\n```\n12 reps · Move through 70° at the shoulder · Hold for 2 seconds\n```\n\nand progress could show:\n\n```\n0 / 12 valid reps\n```\n\nIf a level has no hold requirement, the UI does not invent one.\n\nThis gave us one source of truth for both tracking behavior and patient instructions.\n\nOne architectural decision I liked about Mendly was separating computer vision from LLM reasoning.\n\nWe do not need Gemini to inspect the patient's camera feed.\n\nThe CV layer can transform movement into structured information first.\n\nConceptually, the result can look something like:\n\n```\n{\n  \"exercise\": \"arm_flexion\",\n  \"attempts\": 9,\n  \"valid_repetitions\": 7,\n  \"range_of_motion\": 83,\n  \"tracking_quality\": \"good\"\n}\n```\n\nThe exact schema can change.\n\nThe important part is the boundary.\n\nThe computer vision layer answers:\n\n**What happened?**\n\nThe AI planning layer answers:\n\n**How should that performance history influence a future exercise set?**\n\nThis is an important detail about Mendly's architecture.\n\nThe patient's active session follows an exercise set that has already been reviewed and approved.\n\nGemini is not sitting inside the rep counter deciding what exercise the patient should perform next.\n\nThe flow is closer to:\n\n```\nPractitioner defines boundaries\n        ↓\nAI drafts a future exercise set\n        ↓\nDeterministic guardrails validate it\n        ↓\nPractitioner reviews / edits / approves\n        ↓\nPatient performs the approved set\n        ↓\nComputer vision measures performance\n        ↓\nValidated results are stored\n        ↓\nPractitioner later requests another draft\n        ↓\nBackboard / Gemini use the accumulated history\n```\n\nThat distinction was important to us because Mendly is not trying to make the LLM an autonomous therapist.\n\nThe model helps with planning.\n\nThe practitioner stays in control.\n\nOur stack included:\n\nFor the adaptive planning side, Mendly sends a planning request through **Backboard**, which routes it to a Gemini model.\n\nBackboard also gives us patient-specific planning continuity.\n\nBut memory is not treated as the source of truth.\n\nWhen a new draft is requested, Mendly still sends the current practitioner constraints and recent performance information.\n\nWe also kept a deterministic fallback path.\n\nDuring testing, we actually hit a case where our Backboard API key worked, but LLM chat was unavailable because of account credit.\n\nInstead of breaking the entire practitioner workflow, Mendly could fall back to a local rules-based proposal.\n\nThat experience reinforced a principle I want to keep using in future AI projects:\n\n**AI should improve a workflow, not become a single point of failure.**\n\nBecause Mendly deals with rehabilitation, we did not want the model to have unlimited freedom.\n\nThe practitioner can define constraints such as:\n\nThe model drafts within those constraints.\n\nThen deterministic guardrails check the proposal again.\n\nThen the practitioner makes the final decision.\n\nThe simplest way I think about it is:\n\n```\nPractitioner defines the box\n        ↓\nAI proposes inside the box\n        ↓\nCode checks the box\n        ↓\nPractitioner approves the result\n```\n\nThat architecture was much more interesting to me than simply adding a chatbot to the application.\n\nWe also built a prototype safety-support layer around motor exercises.\n\nThe system can surface signals related to things like a rapid downward movement followed by a low body position, remaining low for an extended period, or disappearing from tracking for a long time.\n\nThose events can create alerts for the practitioner.\n\nBut this is something I would describe carefully.\n\nMendly is **not** a clinically validated fall detector.\n\nIt is a prototype safety-support mechanism intended to surface potentially concerning situations to a human.\n\nThat distinction matters, especially in a healthcare-oriented project.\n\nBefore Mendly, it was easy for me to think of computer vision as:\n\n```\ninput\n  ↓\nmodel\n  ↓\nprediction\n```\n\nAfter spending the hackathon debugging the movement pipeline, I started thinking about the whole system instead:\n\n```\ninput\n  ↓\nprediction\n  ↓\nconfidence\n  ↓\nvalidation\n  ↓\ntemporal reasoning\n  ↓\nstructured data\n  ↓\napplication behavior\n```\n\nThe model is only one component.\n\nThe engineering around the model determines whether its output is actually useful.\n\nA system that sometimes says:\n\n\"I don't know\"\n\ncan be much safer and more useful than one that always produces a number.\n\nThat was probably the biggest lesson I took away from working on Mendly.\n\nBy the end of HopHacks, Mendly had become a working prototype connecting:\n\n```\nPractitioner-defined constraints\n        ↓\nAI-assisted planning\n        ↓\nPractitioner approval\n        ↓\nPatient exercise session\n        ↓\nCamera-based movement measurement\n        ↓\nValidated performance history\n        ↓\nFuture adaptive planning\n```\n\nOur team finished as a **Top 11 finalist out of 87 projects** and won **Best Use of DigitalOcean**.\n\nBut the result I care about most is that I left the hackathon thinking differently about AI systems.\n\nThe interesting part was not simply using MediaPipe or Gemini.\n\nIt was connecting the physical world to an intelligent system without pretending that noisy observations were perfect.\n\nFor me, the core problem became:\n\n**How do you turn imperfect observations into reliable information that an intelligent system can reason over?**\n\nThat problem applies far beyond rehabilitation.\n\nAnd it is probably the part of Mendly I enjoyed building the most.\n\nMendly started as a hackathon project, but building its computer vision pipeline changed how I think about AI engineering.\n\nIt is easy to focus on making models smarter.\n\nBut when AI interacts with real-world data, the quality of the reasoning depends heavily on the quality of the information we give it.\n\nSometimes the most important engineering decision is not:\n\n\"What should the model predict?\"\n\nbut:\n\n**\"Do we actually know enough to make a prediction at all?\"**\n\nThat was the lesson behind a lot of the fixes we made during HopHacks.\n\nAnd it is one I plan to carry into the next system I build.", "url": "https://wpnews.pro/news/how-our-top-11-hophacks-project-turned-noisy-mediapipe-pose-data-into-reliable", "canonical_source": "https://dev.to/jasonpg/how-our-top-11-hophacks-project-turned-noisy-mediapipe-pose-data-into-reliable-rehabilitation-411d", "published_at": "2026-09-30 21:42:46+00:00", "updated_at": "2026-09-30 21:46:39.421857+00:00", "lang": "en", "topics": ["computer-vision", "artificial-intelligence", "ai-products", "ai-agents"], "entities": ["Mendly", "MediaPipe", "HopHacks", "DigitalOcean"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-our-top-11-hophacks-project-turned-noisy-mediapipe-pose-data-into-reliable", "markdown": "https://wpnews.pro/news/how-our-top-11-hophacks-project-turned-noisy-mediapipe-pose-data-into-reliable.md", "text": "https://wpnews.pro/news/how-our-top-11-hophacks-project-turned-noisy-mediapipe-pose-data-into-reliable.txt", "jsonld": "https://wpnews.pro/news/how-our-top-11-hophacks-project-turned-noisy-mediapipe-pose-data-into-reliable.jsonld"}}