Structure-from-Motion (SfM)
Imagine walking around a statue while taking photographs. Although each image is flat, the changing viewpoint reveals which features are near, far, and seen from different angles. Structure-from-Motion (SfM) uses that same idea to reconstruct a scene’s 3D shape and infer where each camera was when it captured an image.
How SfM builds a sceneSfM starts with a collection of overlapping photographs or video frames. It finds distinctive image features—such as corners, textured patches, or logo details—using methods such as SIFT. It then matches features believed to represent the same real-world point across images. From how those matches shift between views, SfM estimates each camera’s pose: its position and orientation. Once camera poses are known well enough, matched rays from multiple cameras can be intersected through triangulation, producing a sparse 3D point cloud. Finally, bundle adjustment jointly refines camera poses and 3D point locations to minimize reprojection error: the gap between predicted and observed feature locations in the images.
What it produces and why it mattersThe result is usually a sparse geometric reconstruction, plus calibrated camera poses. SfM is a foundation for:
- Creating 3D models of buildings, archaeological sites, or products from ordinary photographs.
- Estimating a moving camera’s path in robotics, drones, and autonomous-vehicle perception.
- Supplying camera poses for dense reconstruction methods, including multi-view stereo and neural radiance fields.
- Measuring and inspecting objects in industrial image-capture setups.
SfM does not directly know absolute scale from ordinary monocular images: a small object photographed nearby can resemble a large object photographed farther away. GPS, known distances, stereo cameras, or depth sensors can resolve that ambiguity. It also relies on reliable overlap and visual texture; blank walls, reflections, repeated patterns, and moving people create incorrect matches. COLMAP is a widely used practical SfM system that automates feature matching, reconstruction, and bundle adjustment.
Structure-from-Motion (SfM) is a computer-vision technique that reconstructs a scene’s 3D structure and estimates camera poses from overlapping 2D images taken from different viewpoints. By matching visual features across images and solving their geometry, SfM produces sparse point clouds and calibrated camera trajectories. It is fundamental for photogrammetry, 3D reconstruction, augmented reality, and initializing downstream mapping or neural rendering systems.
Imagine walking around a statue and taking photos from different angles. Even without measuring it, you can compare the pictures and build a sense of the statue’s shape and where you stood for each photo. Structure-from-Motion (SfM) gives computers that same ability.
It takes a collection of overlapping photos or video frames and uses the changing viewpoints to infer a rough 3D scene: the shape and positions of objects, plus the camera’s path through the space. This matters because ordinary images are flat. SfM helps turn them into useful 3D maps for drone surveys, virtual tours, archaeology, film effects, and robots navigating unfamiliar places.