MediaPipe Pose 33 Landmarks: The Complete Guide — Index List and Measured Behavior the Docs Don't Cover

Published: May 2, 2026 (Updated: October 5, 2026)
For: Developers using MediaPipe Pose Landmarker / anyone implementing pose estimation

This article is for you if:

  • You want the index and body part of each of the 33 landmarks
  • You can't decide between the lite, full, and heavy models
  • visibility comes back empty, the depth z looks wrong, or people in your photos aren't detected

Introduction

My 3D drawing-mannequin app, PoseMirror, takes a single photo, gets 33 points from MediaPipe Pose Landmarker, and turns them into a pose on a 3D model.

What kept tripping me up while building it wasn't misremembering indices. It was behavior the official docs don't mention. Code I wrote to use visibility was reading undefined every time. An arm that is perfectly straight in the photo bent at the elbow on the 3D model. Photos without a visible face didn't detect the person at all.

This guide reorganizes what the official docs already cover (the 33-point list, what the outputs mean, the three models) and adds the undocumented behavior, measured on 120 people photos from Unsplash and shown as figures.

The 33 landmarks (official)

Full-body layout

Official layout of the 33 MediaPipe Pose landmarks: a front-facing stick figure with indices 0 to 32 at its joints. The body part for each index is in the tables below

Source: Pose landmark detection guide (Google, CC BY 4.0)

Each hand has four points — wrist, pinky, index, and thumb — and the wrist, pinky, and index form a triangle.

Face and head (0–10)

Index Official name
0 nose
1 left eye (inner)
2 left eye
3 left eye (outer)
4 right eye (inner)
5 right eye
6 right eye (outer)
7 left ear
8 right ear
9 mouth (left)
10 mouth (right)

Upper body (11–22)

Index Official name
11 left shoulder
12 right shoulder
13 left elbow
14 right elbow
15 left wrist
16 right wrist
17 left pinky
18 right pinky
19 left index
20 right index
21 left thumb
22 right thumb

Points 17–22 are only enough to tell which way the hand is facing. If you need finger bending, use the dedicated Hand Landmarker (21 points per hand) alongside it.

Lower body (23–32)

Index Official name
23 left hip
24 right hip
25 left knee
26 right knee
27 left ankle
28 right ankle
29 left heel
30 right heel
31 left foot index
32 right foot index

What the output contains (official)

Pose Landmarker returns two sets of coordinates per person.

Output x, y z Origin and units
landmarks Normalized to 0–1 by image width and height Depth. Smaller values are closer to the camera. Roughly the same scale as x Only z uses the midpoint of the hips as origin
worldLandmarks Real-world 3D coordinates Same as left Origin at the midpoint of the hips, in meters

Each point carries visibility (the likelihood that the point is visible in the image — inside the frame and not occluded). The official docs also list presence, the detection confidence, as an output.

The main options:

Option (Web name) Default Meaning
runningMode IMAGE IMAGE / VIDEO / LIVE_STREAM
numPoses 1 Maximum number of people to detect
minPoseDetectionConfidence 0.5 Threshold for person detection
minPosePresenceConfidence 0.5 Threshold for treating a pose as present
minTrackingConfidence 0.5 Threshold for keeping tracking from the previous frame in video
outputSegmentationMasks false Whether to also output a person segmentation mask

Processing runs in two stages. A person detector first estimates the midpoint of the hips, the radius of a circle enclosing the whole body, and the tilt of the line joining shoulders and hips, which locates the person. That region is passed to the landmark model. In video, the detector only runs on the first frame and when the person is lost; every other frame derives the region from the previous frame's landmarks.

Code example (Tasks API, browser)

import { PoseLandmarker, FilesetResolver } from "https://cdn.jsdelivr.net/npm/@mediapipe/tasks-vision@1.0.1";

const vision = await FilesetResolver.forVisionTasks(
  "https://cdn.jsdelivr.net/npm/@mediapipe/tasks-vision@1.0.1/wasm"
);
const poseLandmarker = await PoseLandmarker.createFromOptions(vision, {
  baseOptions: {
    modelAssetPath:
      "https://storage.googleapis.com/mediapipe-models/pose_landmarker/pose_landmarker_heavy/float16/1/pose_landmarker_heavy.task",
    delegate: "GPU",
  },
  runningMode: "IMAGE",
  numPoses: 1,
});

const result = poseLandmarker.detect(imageElement);
const lm = result.landmarks[0];        // 33 points relative to the image
const world = result.worldLandmarks[0]; // 33 points in meters
if (lm) {
  const leftShoulder = lm[11]; // { x, y, z, visibility }
  console.log(leftShoulder.x, leftShoulder.y, leftShoulder.visibility);
}

Plenty of articles still use the legacy @mediapipe/pose package (receiving results via pose.onResults(...)). For new code, use the Tasks API above.

The three models: lite / full / heavy (official)

Model File size* Accuracy PCK@0.2 (yoga / dance / HIIT) Latency Pixel 3 / MacBook Pro 2017
lite 5.8MB 90.2% / 92.5% / 93.5% 20ms / 25ms
full 9.4MB 95.5% / 96.3% / 95.7% 25ms / 27ms
heavy 30.7MB 96.4% / 97.2% / 97.5% 53ms / 38ms

*File sizes aren't in the docs; I measured the .task files at the download URLs in October 2026. Accuracy and latency come from the legacy MediaPipe Pose documentation.

All three take the same input sizes: 224×224 for the person detector and 256×256 for the landmark model.

What the docs don't cover (measured)

The results below come from running 120 people photos from Unsplash (dance, yoga, sports, sitting, crouching, and so on) through all three models.

  • Environment: Windows 11 / Chromium (Playwright) / GeForce RTX 4070 / Core i7-13700F
  • Library: @mediapipe/tasks-vision 1.0.1 (several versions compared for the visibility section only)

1. Older versions don't output visibility

Even if your code is written to use visibility, an older browser version gives you undefined. I checked the keys on each point with the heavy model across versions.

@mediapipe/tasks-vision version Keys per point
0.10.0 – 0.10.10 x, y, z only
0.10.11 – 1.0.1 x, y, z, visibility

The presence value listed in the docs was not attached to individual points, even in 1.0.1.

A check like lm.visibility > 0.5 treats every joint as hidden on older versions, because undefined > 0.5 is false. Code that fills in a default with lm.visibility ?? 1 does the opposite and treats every joint as visible. Neither throws an error, which makes the bug easy to miss.

2. lite vs. heavy shows up in which skeletons come out right

The same two photos with lite, full, and heavy skeletons overlaid. In the first row (yoga, upward-facing dog), lite and full lose the legs and the skeleton collapses onto the arms; only heavy follows the legs stretched along the floor. In the second row (backlit lunge), lite's legs fall apart while full and heavy get both the front and back leg. Latencies: lite 30ms and 26ms, full 33ms and 28ms, heavy 43ms and 36ms

In the yoga photo on the first row, lite and full lose the legs and the skeleton collapses onto the arms. Only heavy gets the legs. In the backlit lunge on the second row, only lite's legs fall apart.

Across all 120 photos:

Model Photos detected Latency (median)
lite 88 / 120 31ms
full 85 / 120 33ms
heavy 90 / 120 42ms

The number of detected photos barely changes. The difference is whether the skeleton is right on the photos that were detected. As in the figure, lite tends to fail by finding the person but putting the legs in the wrong place.

PoseMirror analyzes one photo at a time; it doesn't have to process dozens of frames per second like video. So I chose heavy for correct skeletons over a 10ms gain. The catch is file size: at 30.7MB it's about five times lite, and the first load is slower for it.

3. CPU and GPU detect different photos

Setting baseOptions.delegate to "CPU" or "GPU" changed which photos were detected at all.

Model Detected on CPU only Detected on GPU only
lite 10 photos 4 photos
full 10 photos 5 photos
heavy 7 photos 7 photos

With heavy, CPU and GPU both detected 90 photos, but 7 on each side were different photos. Running GPU twice, the detected/not-detected result matched on all 120 photos. Since the same environment gives the same result every time, I take this to be a difference in how CPU and GPU compute, not random noise (I haven't confirmed the root cause).

A photo that works on your machine may fail in a user's environment. When reproducing a bug report, match the delegate first.

4. Using worldLandmarks depth as-is throws off the elbows

Left: a photo of a woman dancing, with the arm reaching to the lower right fully straight. Top right: worldLandmarks applied as-is to the PoseMirror 3D mannequin — the same arm bends at the elbow and the hand drops near the hip. Bottom right: with z scaled to 1/3, the arm stays straight as in the photo. Each result is shown from the front and from the side

In the photo, the arm reaching to the lower right is straight. Using worldLandmarks as-is on the 3D model bent that arm at the elbow and dropped the hand near the hip (top right).

The error comes from the depth z. Across 20 poses built from photos in PoseMirror, elbow angles were off from the photo by +34.1° on average, and all 20 were off by more than 15°. x and y are the positions in the image itself, but z can only be inferred from a single photo, which I believe is why its error is larger.

So I convert the x and y of landmarks back to pixels and scale only z down to 1/3 before passing it to the 3D model. On the same 20 poses, the average elbow error dropped to +10.3° (bottom right). The 1/3 factor matches the default (pose_model_z_depth_scale = 3) of XR Animator, a motion-capture app for VTubers.

// Back to image pixels, then scale only z to 1/3 (z shares x's scale, so multiply by width)
const px = lm.map(p => [p.x * imgW, p.y * imgH, (p.z * imgW) / 3]);

The direction of the error depends on the photo. Among the 120 photos, some straight arms bent (like the one in the figure) and some bent arms straightened. Measure on photos from your own use case to pick the scale factor.

5. Half the photos with no person found worked once given a region

Four photos run through PoseMirror. A seated person shown only from the neck down and an upside-down yoga pose were recovered by the fallback detection. A runner with motion-blurred legs and a person lying face down on a mat were not found

In 28 of the 120 photos, no model found a person. Lined up, the common cases are photos showing only the body below the neck, upside-down poses, people small in the frame, and motion blur.

When a clearly visible person still isn't detected, it's because the first-stage person detector drops it. The second-stage landmark model returns a skeleton as long as it's given a region. So when MediaPipe can't find a person, PoseMirror gets the body position from MoveNet (17 points), another Google model, and passes that region directly to MediaPipe's landmark model.

This fallback recovered 14 of the 28 photos (9 via MediaPipe's landmark model, 5 by substituting MoveNet's 17 points). The two on the left of the figure are examples. Blurred photos and the person lying face down on the mat stayed undetected even with the fallback.

6. Hidden joints still get coordinates, but lower visibility

Two photos with each joint colored by visibility: green is visible, red is hidden. For the dancer on the left, the wrist behind her head and the far arm hidden behind her body are red. For the crouching person on the right, the ankles and heels hidden behind the knees are red or orange

With a version that outputs visibility (0.10.11 and later), I colored each joint by its value. Green joints are visible; red joints are hidden.

For the dancer on the left, the wrist behind her head and the far arm in the shadow of her body are red. For the crouching person on the right, the ankles and heels hidden behind the knees are red. Hidden joints still come back with coordinates, but the positions are guesses. Check visibility to decide whether to use those coordinates or fill them in some other way.

The points PoseMirror uses

PoseMirror's 3D model uses 23 of the 33 points.

Purpose Points
Head direction 0 (nose), 7, 8 (ears)
Arms 11–16 (shoulders, elbows, wrists)
Hand direction 17–20 (pinky and index)
Torso and legs 23–28 (hips, knees, ankles)
Foot direction 29–32 (heels, toes)

The eyes (1–6) and mouth (9, 10) aren't used, because the nose and ears already set the head direction. The thumbs (21, 22) aren't used either, since pinky and index are enough for the hand direction.

Summary

The official diagram and tables are all you need for the indices. The time goes into the behavior around them. When you build something new, checking things in this order saves rework.

First, check your @mediapipe/tasks-vision version. At 0.10.10 or earlier there's no visibility, so any check on it silently does nothing. Next, decide on the delegate and test locally with the same setting, because CPU and GPU detect different photos. If you're driving a 3D model, don't use worldLandmarks depth as-is; compare against the photo to decide how much to scale z. And if many photos come back with no person, adding a fallback that passes a region from another model recovers about half of them.

You can try everything in this article by dropping a single photo into PoseMirror.

Watch 33 landmarks turn into a 3D model.

Drop in one photo; the heavy model analyzes it and poses the 3D mannequin.

▶ Pose a 3D mannequin from a photo

Related (Japanese):

Related (English):

References

Photos: From Unsplash. Mathilde Langevin, Matea Brajdić, Morgan Petroski, syahmi syahir, Wesley Tingey, Pierre-Antoine FRANCK, Alina Rubo, Chris Yang, SHAYAN Rostami

← Back to portfolio TOP