Notes

Point Cloud

A point cloud is a collection of dots that describes the shape of something in three dimensions. Instead of representing a scene as a grid of colored pixels, it records where measured points sit in space—like a digital scatter of tiny survey markers around a car, room, face, or factory part.

What each point contains
Each point has at least three coordinates: (x, y, z), which specify its position along width, height, and depth. Points can also carry extra measurements, such as color (RGB), brightness or laser-return strength (intensity), a surface direction (normal), or a timestamp. A LiDAR sensor on an autonomous vehicle, for example, sends out laser pulses and builds a point cloud from their return times. Depth cameras, stereo cameras, and photogrammetry systems can produce them too.

Why point clouds need special handling
Unlike an image, a point cloud has no fixed pixel grid or natural left-to-right ordering. It can be sparse in distant areas, dense nearby, and partly missing where one object blocks another. Vision systems therefore group nearby points, estimate surfaces, and identify geometric patterns. Common operations include:

  • Voxel downsampling, which reduces point count by keeping representative points in small 3D cells.
  • Segmentation, which separates road, vehicles, pedestrians, walls, or individual objects.
  • Registration, which aligns scans taken from different viewpoints; ICP (Iterative Closest Point) is a classic method.

Where they matter
Point clouds give computer vision direct geometric evidence: how far away an object is, how large it is, and what 3D shape it has. This supports obstacle detection for self-driving vehicles, robot grasping, 3D scans for medical planning, and inspection of manufactured parts against a reference model. Libraries such as Open3D and the Point Cloud Library (PCL) provide practical tools for loading, filtering, visualizing, and aligning point clouds.

A point cloud is a set of points in 3D space, typically represented by x, y, z coordinates and optional attributes such as color, intensity, or surface normals. It is commonly produced by LiDAR, depth cameras, or multi-view reconstruction. Point clouds provide a direct representation of scene geometry, enabling 3D detection, mapping, localization, reconstruction, and robotic navigation.

Imagine mapping a room by placing a tiny dot wherever you notice a real surface: one dot on the chair, thousands on the floor, and many more along the walls. From far enough away, that cloud of dots starts to look like the room itself.

A point cloud is this kind of 3D dot map. Each point marks a location in the real world, often with extra details such as colour. Cameras with depth sensing, laser scanners, robots, and self-driving cars use point clouds to understand the shape and position of nearby objects. They matter because the real world is three-dimensional: knowing a car is present is useful, but knowing exactly where it is in space is safer.