writing archive
May 25, 20268 min read

Building Who Do I Run Like

A running-form similarity website that starts as a fun upload-and-match toy, then immediately becomes a computer vision pipeline problem.

computer-visionrunningmachine-learningproduct

The idea is simple enough that it sounds like it should already exist.

Upload a short clip of yourself running. The site analyzes your form. A few seconds later it tells you which elite runner you move like.

Not who you are. Not whether your mechanics are good. Not whether you are injured, efficient, overstriding, undertraining, or secretly one tempo run away from greatness. Just the fun version: who do you run like?

That was the product line I wanted to hold onto. Entertainment-first. Form resemblance, not coaching. The second you start using computer vision on bodies, the wording matters. A toy can turn into a fake diagnostic tool very quickly if you let the copy get too confident.

So I built it around a narrower promise: compare a short running clip against a reference library of elite runners and return a form match.

Who Do I Run Like landing page mockup with upload card, runner comparison reveal, and athlete cards.
Who Do I Run Like landing page mockup with upload card, runner comparison reveal, and athlete cards.

The Website Part

The public site is the easy-to-understand surface area. Big headline. One running figure split between "you" and an elite athlete. Upload a short clip. Compare against the world's best.

The first prototype is a Next.js landing page with three main pieces:

  • a hero comparison reveal where the pointer wipes between a gray runner and Jakob Ingebrigtsen
  • an upload card that accepts MP4, MOV, or WebM clips under 20MB
  • an athlete strip showing the kind of reference library the product is building toward

The UI is intentionally more playful than clinical. It should feel closer to a "which pro do you play like?" sports internet thing than a movement lab report. That choice drives the technical side too. I do not need a perfect biomechanical model to make the product interesting. I need a defensible resemblance pipeline, clear confidence language, and enough visual evidence that the result feels earned.

The Part That Got Complicated

The hard part is not making a page where someone can upload a video. The hard part is answering the question after they do.

The first corpus target is 30 elite distance runners, split across three buckets:

  • 800m / 1500m
  • 5k / 10k
  • marathon / road

Each runner needs 3-5 approved reference segments. That sounds small until you try to build it. You do not just need "a video of Eliud Kipchoge." You need a short segment where the target runner is visible, not buried in a pack, moving in a useful camera angle, with enough frames to capture rhythm, posture, arm swing, knee recovery, and foot path.

So the repo has two halves. There is the public-facing site in site/, and there is the offline CV pipeline that builds and evaluates the internal reference library.

The ingestion loop starts with YouTube discovery, turns candidates into review queues, scores whether the footage probably contains usable full-body running, and then sends the best clips through a local review UI. From there, a human still chooses the useful window and identifies the target runner. That manual step is not a failure. It is the thing that keeps the rest of the pipeline honest.

The Pipeline Shape

The pipeline I am building looks roughly like this:

Text
discover candidate videosscore likely running footagereview and approve short windowsselect the target runner oncetrack that runner through the clipgenerate a runner maskextract pose landmarksoptionally add DensePose body mapsfuse confidence signalscompile form featurescompare against reference segmentsreturn top form matches

A Real Clip Through The Pipeline

Here is a short Cole Hocker reference segment from the 2024 U.S. Olympic Trials. This is what an approved running clip looks like after the review step: a few seconds, one target runner, enough full-body motion to extract stride rhythm and upper-body mechanics, and a cleaner camera angle than most pack-running footage gives you.

This is also why the project cannot just say "find the runner" and move on. Hocker is visible and relatively cleanly framed, but there are still other athletes nearby, the camera is moving, and the target runner changes scale as the shot develops. Even a clean clip still needs identity logic.

The next artifact is the masked runner clip. The point is not to make a cool blue silhouette, though it does look pleasantly sci-fi. The point is to isolate the target runner so later stages are not distracted by the rail, the background, or the athlete next to him.

The mask is useful evidence, but it still is not the matching signal. This is the part I care about for comparison: the pose sequence. Once the clip has been reduced to landmarks over time, the runner's kit, skin tone, background, broadcast graphics, and lighting are mostly gone. What remains is posture, limb timing, arm carriage, knee recovery, hip motion, and stride rhythm.

The fused overlay is the QA view. It puts the pieces back together: the target runner, the mask, the body-region evidence, and the pose. I use this to inspect whether the pipeline is trusting the right frames and joints. If the overlay looks wrong, the matcher does not get to pretend the arrays are fine.

For this clip, the feature compiler ends up with 120 usable frames out of 120. The fused confidence mean is about 0.91 and the pose visibility mean is about 0.85, which is why this is a better demo clip. It gives the matcher a cleaner stride sample without needing to explain away a bunch of questionable frames first.

The important decision is that pose sequence is the matching truth.

The runner mask and DensePose output are useful, but they are not the main answer. The mask isolates the target runner from the rest of the frame. DensePose can tell me which visible pixels correspond to body regions. Both are great for QA and confidence weighting. But the thing I actually want to compare is motion over time: landmarks, joint angles, bone vectors, torso axis, shoulder and hip carriage, arm rhythm, lower-leg recovery path, and stride timing.

That distinction matters because DensePose is visually seductive. It looks rich and precise. But it is also heavier, slower, more sensitive to camera angle and visible surface area, and easier to misuse as appearance similarity. A runner's kit, body shape, lighting, and camera crop can start leaking into the result if you let the pixel representation dominate.

Pose is cheaper, more interpretable, and closer to the product claim. If the site says "you run like Jakob," the result should come from how the body moves, not from whether the clip happened to have a similar silhouette.

Identity Is Its Own Problem

Another early trap: segmentation is not identity.

A model can produce a clean person mask and still be tracking the wrong person. That matters in running footage because athletes overlap constantly. Packs compress. A camera pans. Someone passes in front of the target runner. If the mask silently switches from one runner to another, every downstream feature is poisoned.

So the current architecture treats identity explicitly. The reviewer selects the target runner once with a click or loose box. Then detector and tracker outputs preserve that identity through the clip. ReID signals help decide whether the same target survived occlusion or whether the pipeline should mark an identity risk. SAM 3.1 mask generation runs after that, gated by the chosen target track.

That ordering is the whole game:

Text
identity firstmask secondpose thirdmatch last

If you flip those around, the demo can still look cool, but the data stops meaning what you think it means.

The Form Artifact

The current feature compiler writes two artifacts:

  • form_features.json for metadata, quality summaries, and explainable features
  • form_features.npz for the dense arrays used in sequence matching

The JSON side stores things like frame count, usable frame rate, camera angle bucket, average confidence, pose visibility, torso lean, arm swing amplitude, knee lift proxy, vertical oscillation proxy, and stride rhythm proxy.

The array side stores the real comparison material: normalized pose coordinates, world landmarks when available, joint weights, frame weights, bone vectors, joint angles, angular velocity, DensePose region coverage, valid-frame masks, and timestamps.

For the first matcher, I am staying deliberately boring: exact weighted sequence matching before adding vector search or learned embeddings. The corpus is small enough that I do not need FAISS yet. More importantly, exact matching makes the mistakes easier to inspect. When a match is wrong, I want to know whether it came from bad landmarks, bad camera-angle bucketing, bad confidence weights, or a genuinely weak comparison rule.

Embeddings can come later. They are an experiment, not the foundation.

Confidence Language

The product also avoids percentage matches for now.

"You are an 87% match with Mo Farah" sounds precise in a way the system does not deserve yet. It is also less interesting than an explanation you can actually understand.

The target output is more like:

  • top form match
  • top three alternative runners
  • the matched reference segment for each result
  • confidence as high, medium, or low
  • explanation tags like "similar compact arm swing," "similar torso lean," or "similar stride rhythm"
  • side-by-side playback of the query clip and matched reference segment

That is enough to make the experience feel concrete without pretending this is a lab-grade biomechanics assessment.

Why This Is Fun

This project sits in a nice little tension.

The user-facing idea is unserious in the best way. You upload a clip and get told you look like Faith Kipyegon, Cole Hocker, Sifan Hassan, or Grant Fisher. It is made for curiosity, group chats, and runners who already overthink every frame of their stride anyway.

Underneath that, the engineering is very real. Video curation. Identity tracking. ReID. SAM 3.1 masks. Pose extraction. DensePose as a confidence layer. Feature contracts. QC metrics. Active learning queues. Sequence similarity.

That is the kind of project I like: a clean product question sitting on top of a messy data problem.

The next milestone is to finish the local reference feature library and get the first pose-sequence matcher returning runner-level form matches. After that comes explanation tags and side-by-side playback. The site already knows how to promise the experience. Now the pipeline has to earn it.