# How our Top 11 HopHacks project turned noisy MediaPipe pose data into reliable rehabilitation metrics for adaptive AI planning

> Source: <https://dev.to/jasonpg/how-our-top-11-hophacks-project-turned-noisy-mediapipe-pose-data-into-reliable-rehabilitation-411d>
> Published: 2026-09-30 21:42:46+00:00

What if a rehabilitation exercise could understand **how** a patient moved, instead of only asking whether they finished?

That was one of the questions behind **Mendly**, an AI-powered stroke rehabilitation platform my team built during HopHacks.

By the end of the hackathon, Mendly became a **Top 11 finalist out of 87 projects** and won **Best Use of DigitalOcean**.

I worked mainly on the computer vision side of the project: pose tracking, joint-angle and range-of-motion measurement, valid-repetition detection, tracking-loss handling, camera reliability, and the layer that turns movement into structured performance data for our adaptive AI pipeline.

The part I found most interesting was not simply getting MediaPipe to recognize a person.

It was figuring out how to turn imperfect camera observations into data that the rest of the system could actually trust.

The pipeline looked roughly like this:

```
Camera
  ↓
Pose Detection
  ↓
Joint Angles / Range of Motion
  ↓
Reliability Checks
  ↓
Repetition Detection
  ↓
Structured Performance Data
  ↓
Future Adaptive Planning
```

That sounds straightforward on paper.

It wasn't.

We used **MediaPipe Pose** to estimate landmarks on the patient's body.

Instead of trying to reason directly about image pixels, we could work with points representing joints such as the shoulder, elbow, wrist, hip, knee, and ankle.

For example, if we know the positions of the shoulder, elbow, and wrist, we can calculate the angle at the elbow.

Given two vectors around a joint:

```
v₁ = Shoulder - Elbow
v₂ = Wrist - Elbow
```

we can calculate the angle between them with the dot product:

```
θ = arccos((v₁ · v₂) / (|v₁||v₂|))
```

Once we have joint angles over time, we can start measuring things like range of motion.

If an elbow moves from roughly:

```
155° → 72°
```

then the observed range is approximately:

```
83°
```

This is already much more useful than simply knowing that "the arm moved."

But getting an angle is the easy part.

The harder question is whether that angle should be trusted.

This was probably the biggest thing I learned from the project.

At first, it is easy to think of pose estimation like this:

```
MediaPipe detected a wrist
        ↓
therefore the wrist position is correct
```

Real camera input is much messier.

A patient can move partly outside the frame. A joint can become occluded. Lighting can change. Landmarks can jump. The patient can turn sideways. The camera can be positioned badly.

A single bad landmark can propagate through everything after it:

```
bad landmark
    ↓
bad angle
    ↓
bad repetition
    ↓
bad performance data
    ↓
bad future planning context
```

That made reliability part of the actual product, not just an implementation detail.

Some of the most interesting bugs appeared when we asked a simple question:

**What should happen when the camera does not know?**

Suppose the tracker sees:

```
Frame 1: 40°
Frame 2: 55°
Frame 3: tracking fails
```

What should Frame 3 become?

There are two tempting answers.

One is:

```
0°
```

The other is:

```
55°
```

Neither is correct.

If tracking failed, we do not know where the patient's arm actually was.

Treating the missing value as zero invents a movement that never happened.

Keeping the previous angle is not much better. The patient may have continued moving while the camera lost them.

So we changed the pipeline to represent failed inference as an **invalid observation**.

```
Pose succeeds
→ process the measurement

Pose fails
→ mark observation invalid
→ do not update movement state
→ do not manufacture an angle
→ do not count a repetition
```

The rule became:

**Unknown is not zero. Unknown stays unknown.**

That sounds obvious after the fact, but it fixed an important class of false movement data.

Another issue was caused by normalized coordinates.

MediaPipe gives landmark positions as normalized X and Y values.

But X is normalized relative to the frame width, while Y is normalized relative to the frame height.

On a 640×480 camera:

```
Δx = 0.1 → 64 pixels
Δy = 0.1 → 48 pixels
```

Those values are both `0.1` in normalized space, but they do not represent the same physical distance in the image.

If we calculate geometry as if the normalized X and Y scales are identical, joint angles can become distorted.

We fixed this by accounting for the actual camera aspect ratio before calculating angles and ray lengths.

At the same time, we kept the original normalized landmarks for things like checking whether a joint was still inside the frame.

This was a good example of a bug that was mathematically small but affected everything built on top of it.

Another thing that looked simple at first was repetition counting.

A naive approach might say:

Count a rep whenever the joint angle crosses a threshold.

That works until real movement starts looking like this:

```
155°
148°
139°
121°
97°
76°
81°
95°
119°
142°
151°
```

Human motion does not happen in perfectly clean steps.

The angle may hover around a threshold, move backward briefly, or contain noise.

So instead of treating each frame independently, we used a state-based approach.

A complete repetition is closer to:

```
Start
  ↓
Movement begins
  ↓
Required range reached
  ↓
Hold requirement satisfied, if needed
  ↓
Movement reverses
  ↓
Return condition reached
  ↓
Valid repetition
```

Crossing one threshold is not automatically a repetition.

A rep has a beginning, progression, target, and completion.

During testing, we found another behavior that looked reasonable in code but felt completely wrong from the patient's perspective.

Imagine the exercise target is:

```
5 valid reps
```

and the patient does:

```
valid
failed
failed
valid
failed
failed
failed
```

They have completed only:

```
2 / 5 valid reps
```

But an earlier completion rule could still stop the exercise after enough total attempts.

That meant a patient could fail several repetitions and somehow "finish" the exercise.

We changed that.

Automatic completion now depends on the number of **valid repetitions**, not the total number of attempts.

```
validRepCount >= targetRepCount
```

Failed attempts can still be recorded.

They just do not count toward the target.

So if the prescription says five valid repetitions, the patient actually needs five valid repetitions.

This distinction also became important for range-of-motion statistics.

Suppose a patient performs a repetition while the camera is barely tracking the required landmarks.

We may still want to know that an attempt happened.

But that does not mean its ROM measurement should be treated as reliable.

Earlier logic could fall back to unreliable repetitions when there were no reliable ones available.

We changed that behavior.

A low-confidence attempt can remain part of the exercise history, but unreliable motion should not contaminate trusted ROM statistics.

I started thinking about this as two different questions:

```
Did something happen?
```

and:

```
Do we trust this measurement?
```

Those are not the same thing.

Pose tracking also depends heavily on how the patient positions the camera.

We built a setup checker to decide whether the current view was usable for the exercise.

One issue was that early bad frames could influence the setup state for too long.

Imagine this:

```
bad setup
bad setup
bad setup
bad setup

patient fixes the camera

good
good
good
```

If the checker effectively remembers the entire history, the patient can remain stuck even after fixing the problem.

We changed the logic to focus on a recent rolling window instead.

That way, old bad setup frames do not permanently punish the patient.

We also kept checking the environment after the patient reached the Ready state.

Otherwise this could happen:

```
Patient is visible
→ Ready

Patient leaves the frame
→ still Ready
```

Instead, if the setup becomes unreliable again, Ready can be revoked.

We also added a timeout.

If the pose stream is running but the system cannot establish a reliable camera view after around 25 seconds, it stops waiting forever and shows the patient a retry option.

That turned camera setup from a one-shot gate into something that could actually recover from failure.

Another bug was less about computer vision and more about product consistency.

An exercise level could require something like:

```
12 reps
70° movement target
2-second hold
```

The tracker knew about the hold.

The patient did not.

The UI originally might only say:

Raise your arm out to the side, then lower it slowly.

From the patient's point of view, they could perform the movement correctly according to the instruction and still fail the tracker.

We fixed this by deriving the patient-facing prescription from the same exercise-level configuration used by the tracking logic.

So the interface could show:

```
12 reps · Move through 70° at the shoulder · Hold for 2 seconds
```

and progress could show:

```
0 / 12 valid reps
```

If a level has no hold requirement, the UI does not invent one.

This gave us one source of truth for both tracking behavior and patient instructions.

One architectural decision I liked about Mendly was separating computer vision from LLM reasoning.

We do not need Gemini to inspect the patient's camera feed.

The CV layer can transform movement into structured information first.

Conceptually, the result can look something like:

```
{
  "exercise": "arm_flexion",
  "attempts": 9,
  "valid_repetitions": 7,
  "range_of_motion": 83,
  "tracking_quality": "good"
}
```

The exact schema can change.

The important part is the boundary.

The computer vision layer answers:

**What happened?**

The AI planning layer answers:

**How should that performance history influence a future exercise set?**

This is an important detail about Mendly's architecture.

The patient's active session follows an exercise set that has already been reviewed and approved.

Gemini is not sitting inside the rep counter deciding what exercise the patient should perform next.

The flow is closer to:

```
Practitioner defines boundaries
        ↓
AI drafts a future exercise set
        ↓
Deterministic guardrails validate it
        ↓
Practitioner reviews / edits / approves
        ↓
Patient performs the approved set
        ↓
Computer vision measures performance
        ↓
Validated results are stored
        ↓
Practitioner later requests another draft
        ↓
Backboard / Gemini use the accumulated history
```

That distinction was important to us because Mendly is not trying to make the LLM an autonomous therapist.

The model helps with planning.

The practitioner stays in control.

Our stack included:

For the adaptive planning side, Mendly sends a planning request through **Backboard**, which routes it to a Gemini model.

Backboard also gives us patient-specific planning continuity.

But memory is not treated as the source of truth.

When a new draft is requested, Mendly still sends the current practitioner constraints and recent performance information.

We also kept a deterministic fallback path.

During testing, we actually hit a case where our Backboard API key worked, but LLM chat was unavailable because of account credit.

Instead of breaking the entire practitioner workflow, Mendly could fall back to a local rules-based proposal.

That experience reinforced a principle I want to keep using in future AI projects:

**AI should improve a workflow, not become a single point of failure.**

Because Mendly deals with rehabilitation, we did not want the model to have unlimited freedom.

The practitioner can define constraints such as:

The model drafts within those constraints.

Then deterministic guardrails check the proposal again.

Then the practitioner makes the final decision.

The simplest way I think about it is:

```
Practitioner defines the box
        ↓
AI proposes inside the box
        ↓
Code checks the box
        ↓
Practitioner approves the result
```

That architecture was much more interesting to me than simply adding a chatbot to the application.

We also built a prototype safety-support layer around motor exercises.

The system can surface signals related to things like a rapid downward movement followed by a low body position, remaining low for an extended period, or disappearing from tracking for a long time.

Those events can create alerts for the practitioner.

But this is something I would describe carefully.

Mendly is **not** a clinically validated fall detector.

It is a prototype safety-support mechanism intended to surface potentially concerning situations to a human.

That distinction matters, especially in a healthcare-oriented project.

Before Mendly, it was easy for me to think of computer vision as:

```
input
  ↓
model
  ↓
prediction
```

After spending the hackathon debugging the movement pipeline, I started thinking about the whole system instead:

```
input
  ↓
prediction
  ↓
confidence
  ↓
validation
  ↓
temporal reasoning
  ↓
structured data
  ↓
application behavior
```

The model is only one component.

The engineering around the model determines whether its output is actually useful.

A system that sometimes says:

"I don't know"

can be much safer and more useful than one that always produces a number.

That was probably the biggest lesson I took away from working on Mendly.

By the end of HopHacks, Mendly had become a working prototype connecting:

```
Practitioner-defined constraints
        ↓
AI-assisted planning
        ↓
Practitioner approval
        ↓
Patient exercise session
        ↓
Camera-based movement measurement
        ↓
Validated performance history
        ↓
Future adaptive planning
```

Our team finished as a **Top 11 finalist out of 87 projects** and won **Best Use of DigitalOcean**.

But the result I care about most is that I left the hackathon thinking differently about AI systems.

The interesting part was not simply using MediaPipe or Gemini.

It was connecting the physical world to an intelligent system without pretending that noisy observations were perfect.

For me, the core problem became:

**How do you turn imperfect observations into reliable information that an intelligent system can reason over?**

That problem applies far beyond rehabilitation.

And it is probably the part of Mendly I enjoyed building the most.

Mendly started as a hackathon project, but building its computer vision pipeline changed how I think about AI engineering.

It is easy to focus on making models smarter.

But when AI interacts with real-world data, the quality of the reasoning depends heavily on the quality of the information we give it.

Sometimes the most important engineering decision is not:

"What should the model predict?"

but:

**"Do we actually know enough to make a prediction at all?"**

That was the lesson behind a lot of the fixes we made during HopHacks.

And it is one I plan to carry into the next system I build.
