Voxel Grid
A voxel grid gives a computer a tidy way to represent a three-dimensional scene. It divides space into many small cubes, much like graph paper divides a page into squares; each tiny cube is called a voxel, short for “volume pixel.”
How the grid represents space
Each voxel covers a fixed region of 3D coordinates: width, height, and depth. Its stored value describes what is inside that region. In a simple occupancy grid, the value says whether the cube is empty or occupied. Richer grids can store:
- Point density from a LiDAR scan
- Color from RGB-D cameras
- Surface distance or a signed distance value
- Learned features used by a neural network
- Semantic labels such as road, car, wall, or bone
Turning point clouds into voxels
A point cloud contains individual 3D measurements, which can be irregularly spaced and awkward for grid-based models. Voxelization assigns every point to a cube based on its coordinates, then combines points in that cube—for example by marking it occupied or averaging their positions and colors. The PCL VoxelGrid filter also uses this idea to downsample dense point clouds: it replaces many nearby points with a representative point. This reduces noise and computation while preserving the scene’s broad shape.
Why resolution matters
Small voxels capture fine details, such as a pedestrian’s outline or a tiny crack in a manufactured part, but require far more memory. Large voxels are efficient but can merge nearby objects or erase narrow structures. This trade-off matters in autonomous driving, where voxel grids convert LiDAR data into inputs for 3D object detectors, and in medical imaging, where CT and MRI scans are naturally stored as voxel volumes for organ or tumor segmentation. Because most 3D space is empty, modern systems commonly use sparse voxel grids, storing only occupied or relevant cells rather than every cube in the volume.
A voxel grid is a three-dimensional array that partitions space into fixed-size volumetric cells, or voxels. Each voxel stores information such as occupancy, density, color, or semantic class, creating a discrete representation of a scene or object. Voxel grids enable spatial reasoning for 3D reconstruction, segmentation, and object detection, but their memory cost grows rapidly as resolution increases.
Imagine dividing a room into thousands of tiny, transparent boxes, like a 3D version of graph paper. Each little box can record whether that bit of space is empty, contains part of a chair, or holds a wall. That collection of boxes is a voxel grid.
A voxel is simply a “pixel with depth.” Pixels describe tiny squares in a photograph; voxels describe tiny cubes in real space. AI systems use voxel grids to make sense of 3D scans from cameras, robots, or medical equipment. They give a machine an organized map of space, helping it recognize objects, avoid obstacles, or inspect structures inside the body.