Cutout
Training images rarely show every object perfectly: a face can be hidden by sunglasses, a car by another vehicle, and a product by glare or packaging. Cutout prepares a vision model for that reality by deliberately hiding a small region of an image during training.
How it worksCutout is a data-augmentation technique that places one or more solid, usually rectangular patches over an input image. The patch is commonly filled with zeros, the dataset’s mean colour, or a random colour. Its position and size are sampled anew for each training image, so the model repeatedly sees different parts missing. The original image remains intact at evaluation time.
- A photograph of a dog might have its ear or background covered.
- An OCR training image might lose part of a character, encouraging recognition from the remaining strokes.
- A defect-inspection model might see a component with an unrelated area masked, rather than learning to rely on a single visual cue.
Without Cutout, a classifier can take shortcuts: it may identify “boat” from a strip of water, or “person” from a face alone. By removing evidence at random, Cutout pushes the network to combine information from multiple regions and develop more robust features. It is conceptually related to dropout, but dropout removes internal neural activations, while Cutout removes pixels in the training image itself.
Practical importanceCutout is especially useful when training data is limited or real deployment includes occlusion, clutter, and partial views. In object detection and segmentation, it must be applied carefully: a mask can hide much of an annotated object while its bounding box or pixel label remains present. Libraries such as Albumentations provide CoarseDropout-style transforms for this purpose. Patch size matters: tiny patches have little effect, while huge patches can erase the evidence needed to learn the label. Good settings create a meaningful challenge without turning images into noise.
Cutout is a data-augmentation technique that randomly masks a rectangular region of a training image, typically by filling it with a constant color or noise. The label remains unchanged, forcing a vision model to learn from multiple visual cues rather than relying on one highly discriminative area. Cutout improves robustness to occlusion and reduces overfitting in image-classification models.
Imagine teaching someone to recognize a dog even when a sofa cushion, a tree, or another dog hides part of it. You would not show only perfect, unobstructed pictures. Cutout does something similar for AI: during training, it covers a small random patch of an image, like placing a sticky note over part of a photo.
This encourages the AI not to rely too heavily on one tiny detail, such as a cat’s ear or a car’s logo. Instead, it learns to use many clues across the image. That matters because real-world photos are often partly blocked, cropped, or messy. Cutout helps vision systems stay reliable when they cannot see everything clearly.