Pose Estimation
Pose estimation teaches a vision system to locate meaningful points on a body, such as shoulders, elbows, knees, and ankles. Rather than merely recognizing that a person is present, it describes how that person is arranged: standing, reaching, crouching, or waving.
How a pose is representedMost 2D pose-estimation models predict a set of keypoints in image coordinates, usually written as (x, y) positions plus a confidence score. The keypoints are connected according to a predefined skeleton: an elbow links to a wrist and shoulder, for example. Modern neural networks commonly produce a heatmap for each joint—a small image in which bright areas indicate likely joint locations. The model selects the strongest location from each heatmap, then groups joints that belong to the same person.
- Single-person pose estimation finds joints for one cropped person.
- Multi-person pose estimation separates and estimates several people in a scene.
- 3D pose estimation estimates depth as well, producing positions such as (x, y, z).
A camera sees only a flat projection of the world, so limbs can overlap, disappear behind objects, or look very different from one viewpoint to another. Loose clothing, motion blur, poor lighting, and crowded scenes add further ambiguity. A raised arm behind someone’s head, for instance, can be difficult to assign to the correct person. Systems such as OpenPose address multi-person scenes with part-confidence maps and limb connections, while libraries such as MediaPipe Pose provide efficient landmark tracking for video.
What pose estimation enablesPose information turns pixels into body movement that other systems can reason about. It supports gesture-controlled interfaces, exercise feedback, sports analysis, animation, and fall detection. In autonomous-vehicle perception, a pedestrian’s posture can help indicate whether they are stepping into the road. In video, tracking keypoints across frames also reveals actions and motion patterns. Without pose estimation, a detector can say “person”; with it, a system gains structured evidence about what that person is doing.
Pose estimation is the task of locating an object’s key points and estimating its spatial configuration in an image or video, such as a person’s joints, body orientation, or a camera’s position. It produces structured geometric information rather than only a class label. Pose estimation enables motion analysis, gesture control, augmented reality, robotics, sports tracking, and human–computer interaction.
Imagine drawing a simple stick figure over a photo of a person: dots at the shoulders, elbows, wrists, hips, knees, and ankles, connected to show how they are standing or moving. Pose estimation teaches a computer to do something like this automatically.
It helps AI identify where key body parts are in an image or video and understand a person’s posture: sitting, waving, running, dancing, or reaching for something. This matters for fitness apps that count exercises, video games that track body movement, sports analysis, animation, and systems that can recognize when someone may have fallen. It is not about knowing who a person is; it is about understanding how their body is positioned.