Notes

Vertical Flip

A vertical flip turns an image upside down: pixels at the top move to the bottom, and pixels at the bottom move to the top. Imagine holding a photograph and rotating it around a horizontal line through its center—not rotating it 180 degrees, but reflecting it across that line.

What changes mathematically
For an image with height H, a pixel at vertical position y is moved to approximately H − 1 − y, while its horizontal position stays unchanged. The image’s width, colors, and pixel values do not change; only their vertical arrangement does. A vertical flip differs from a horizontal flip, which swaps left and right instead. In OpenCV, cv2.flip(image, 0) performs a vertical flip; in PyTorch, torchvision.transforms.RandomVerticalFlip applies it randomly during training.

Why labels must flip too
When vertical flipping is used for training-data augmentation, every associated annotation must undergo the same transformation. This includes:

  • Bounding boxes: their top and bottom coordinates must be remapped.
  • Segmentation masks: the mask image is flipped pixel-for-pixel with the source image.
  • Keypoints: a person’s head, shoulders, hands, and feet need updated vertical coordinates.

Failing to transform labels creates contradictory training examples: the image shows an object in one place while its label claims it is elsewhere.

When it helps—and when it harms
Vertical flips are valuable when “up” and “down” carry little meaning, such as microscope cells, satellite imagery, textured materials, or defect inspection on manufactured parts. They increase visual variety without collecting new images. They are usually a poor choice for ordinary street scenes, faces, handwritten text, and autonomous-driving images: upside-down people, cars, road signs, and letters are unrealistic. A model trained heavily on such examples can spend capacity learning patterns it will never encounter in deployment.

Vertical flip is a geometric image transformation that reflects an image across its horizontal axis, swapping top and bottom pixel rows while preserving left-to-right positions. In computer vision, it is used as data augmentation to increase training diversity and encourage invariance to vertical orientation when such flips preserve label meaning. It is unsuitable for tasks where upright orientation carries essential semantic information.

Imagine turning a photograph upside down: the sky is now at the bottom, and people appear to stand on their heads. That is a vertical flip. It reverses an image from top to bottom while keeping the left and right sides in the same places.

In AI image work, vertical flips are often used to give a model more varied examples during training. For pictures where upside-down versions are still believable—such as microscope images, textures, or some aerial views—this can help the AI focus on the important visual patterns rather than memorising one exact orientation. But it is usually avoided for everyday scenes, where an upside-down car or face would be unrealistic.