Fully Convolutional Network (FCN)
A Fully Convolutional Network (FCN) turns an image classifier into a model that can label every pixel in an image. Rather than answering “what is in this picture?”, it produces a map answering “which pixels belong to a road, person, tumor, sky, or defect?”
How it works
Early convolutional neural networks were designed for classification and ended with fully connected layers, which collapse an image into one fixed-size prediction. An FCN replaces those layers with convolutions, preserving the image’s spatial layout. The network produces a grid of class scores, then enlarges that grid back toward the original image size using upsampling, commonly implemented with transposed convolution. Each output location receives a label such as “background,” “car,” or “building.” Because it contains only convolutional-style operations, an FCN can process images of different dimensions without requiring a fixed input size.
Recovering fine detail
Deep layers recognize high-level concepts well but work at low resolution: they can identify a car while losing its precise outline. The original FCN design addressed this by combining deep, semantic features with earlier, higher-resolution features through skip connections. Its named variants reflect how much detail is restored:
- FCN-32s upsamples a coarse prediction directly.
- FCN-16s combines one earlier feature map.
- FCN-8s combines multiple earlier maps for sharper boundaries.
Why it matters
FCNs made dense, end-to-end pixel prediction practical: a medical scan can separate a lesion from healthy tissue, an autonomous vehicle can mark drivable road area, and a factory camera can outline scratches rather than merely flagging an entire product as faulty. Training compares the predicted class at each pixel with a human-made segmentation mask, usually using pixel-wise cross-entropy. Without spatially preserved predictions, these systems could recognize an object but could not reliably show where it is or trace its boundaries.
A Fully Convolutional Network (FCN) is a neural network that replaces fully connected classification layers with convolutional layers, allowing it to produce spatially aligned, per-pixel predictions for images of varying sizes. In semantic segmentation, an FCN labels every pixel while preserving image layout, typically using upsampling to restore output resolution. It established end-to-end dense prediction as a practical foundation for modern segmentation systems.
Imagine coloring in a map: instead of simply saying “this picture contains a road,” you color every tiny spot that belongs to a road, tree, car, or person. A Fully Convolutional Network (FCN) helps AI do that with images.
It creates a detailed label map, giving each pixel a category. This matters when rough answers are not enough—for example, a self-driving car needs to know exactly where the road ends and the sidewalk begins, while doctors may need to outline a tumor in a scan. FCNs are usually trained with human-labeled examples, so they learn from pictures where the correct regions have already been marked.