HomeFeaturesHow it worksPricingAboutBlogContactGet the app
AI & SCIENCE8 min read

How multi-modal AI reads your squat better than a mirror

A look under the hood at the computer-vision pipeline that turns a 10-second clip into joint-by-joint form feedback — and why combining video, photo and text beats any single signal.

DC
Deepanshu Choudhary
Head of AI · Jun 15, 2026
How multi-modal AI reads your squat better than a mirror
AI squat form analysiscomputer vision fitnessAI personal trainerexercise form checker

Mirrors show you one angle. Your phone shows all of them — and more. When you upload a 10-second clip of your squat to FoodNutrix, a multi-modal AI pipeline breaks it down frame by frame, joint by joint, in ways that no mirror — and very few human coaches — can match at scale.

The anatomy of a bad squat

Most squat errors are invisible to the untrained eye in the moment. Knee cave only appears at the bottom position. Forward lean accumulates across reps. Hip shift is often opposite to where the pain manifests. These compound silently until something pops.

A good strength coach catches these in real time. But most people train alone, or with a buddy who is also mid-set and can't see clearly. That's the gap we built into.

What the AI actually sees

The FoodNutrix video model detects a 17-point body skeleton — one landmark per major joint: ankles, knees, hips, shoulders, elbows, wrists, and key spinal references. It does this across every frame of your video at up to 30 fps.

  • Joint angles: knee valgus, hip crease relative to knee at the bottom, elbow flare on upper-body lifts
  • Temporal patterns: is the descent controlled? Is there a pause at the bottom? Is the drive through the heels?
  • Bilateral symmetry: left/right hip height, weight distribution, shoulder levelness
  • Range of motion: depth relative to parallel, overhead reach on overhead squats

From these signals the model produces a Form Score (0–100) and a ranked list of corrective cues — highest-impact first, so you're never overwhelmed with feedback.

How to read your score

A 94 Form Score means the squat is safe and technically solid. A 72 means there's a correctable fault affecting efficiency or injury risk. A sub-60 means stop and fix before loading more weight.

Combining video with your other data

Multi-modal means the video analysis doesn't live in isolation. The AI knows how heavy you said you went, whether it's a warm-up or working set, how much sleep you got last night, and whether your wearable flagged elevated resting heart rate. A technically imperfect squat on a deload week is a very different finding than the same squat when you're fresh and at your peak.

This context changes the feedback. If you're fatigued the cue becomes 'core bracing breaks down under fatigue — consider ending the session'. If you're fresh it becomes 'brace before you descend and this becomes a 97'.

The limits — and why we're honest about them

Camera angle matters. Side profile gives the richest data. Front-on can detect knee cave but misses forward lean. Oblique angles produce the best bilateral symmetry data. We tell you this in-app and give you a framing guide before you record.

The AI doesn't replace a sports medicine professional. For chronic pain, pathological movement, or post-surgical return-to-training, work with a physio. Our form scoring is designed for technically healthy people training to improve — not for diagnosing injuries.

What comes next

We're working on live real-time feedback — rep-by-rep audio cues while you're mid-set — and expanding the exercise library beyond the current 40 movements. If there's a lift you want scored, let us know at help@foodnutrix.com.

#AIsquatformanalysis#computervisionfitness#AIpersonaltrainer#exerciseformchecker#FoodNutrixvideocoach
DC
Deepanshu Choudhary
Head of AI at FoodNutrix
Try it yourself

Put the science to work — free.

Everything in this article is built into FoodNutrix. No credit card, no timer. Just start.