SLAM
Imagine walking through an unfamiliar building while keeping track of where you are and gradually drawing a floor plan at the same time. SLAM, short for Simultaneous Localization and Mapping, gives a robot, vehicle, or augmented-reality device that same ability: it estimates its own position while building a map of its surroundings.
How SLAM worksA camera captures a stream of images as the device moves. SLAM identifies stable visual details—such as corners, signs, window edges, or textured patches—and tracks how their positions shift between frames. From those shifts, it estimates the camera’s pose: its location and orientation in 3D space. At the same time, it triangulates the observed features to place them in a growing map.
This is a feedback loop: a better map helps estimate the camera’s position more accurately, and a better position estimate makes the map more consistent. Modern visual SLAM systems also perform loop closure. When the camera revisits a previously seen place, the system recognizes it and corrects drift that accumulated during the journey.
Inputs and practical systems- Visual SLAM uses one or more cameras; stereo cameras make depth estimation easier.
- Visual-inertial SLAM combines camera images with an IMU’s accelerometer and gyroscope readings, improving tracking during fast motion or brief visual blur.
- RGB-D SLAM uses color images plus direct depth measurements, such as those from depth cameras.
SLAM enables a robot vacuum to navigate rooms, an autonomous vehicle to align itself with a detailed road map, and an AR headset to keep virtual furniture fixed convincingly on a real floor. It also supports indoor mapping where GPS is unavailable. Without reliable SLAM, small position errors compound: a robot can misjudge obstacles, or virtual objects can visibly slide across the scene. Widely used systems include ORB-SLAM, which detects and matches efficient ORB image features, then refines camera poses and map points together.
SLAM (Simultaneous Localization and Mapping) is the process of estimating a camera or robot’s position while building a consistent map of an unknown environment from sensor data, commonly images, depth, or lidar. SLAM provides the spatial foundation for autonomous navigation, augmented reality, and robotic perception; without accurate localization and mapping, these systems cannot reliably understand or move through 3D space.
Imagine walking through an unfamiliar house in the dark with a notebook. As you move, you sketch where the rooms and furniture are, while also using your sketch to work out where you currently stand. SLAM does this for robots, drones, and self-driving machines.
Short for Simultaneous Localization and Mapping, it helps a machine build a map of its surroundings while figuring out its own position within that map. A robot vacuum, for example, can remember walls and avoid getting lost. SLAM often learns directly from camera or sensor observations rather than needing people to label every object, which connects it to unsupervised learning: finding useful structure in raw data.