Notes

Two-Stream Network

A video is not just a stack of photographs: it also contains movement. A Two-Stream Network is designed to read both parts of that story—what is visible in each frame and how things move from frame to frame.

Two complementary views of video

The classic two-stream design uses two separate neural networks. The spatial stream receives ordinary RGB frames, learning appearance cues such as faces, objects, clothing, scenery, and body poses. The temporal stream receives optical flow: a representation of estimated pixel motion between nearby frames. Optical flow describes direction and speed, usually as colored or multi-channel motion fields. This lets the model distinguish actions that look similar in a single image but differ through motion, such as sitting down versus standing up.

How the streams work together

Each stream produces class predictions or internal feature vectors. The system then fuses them—commonly by averaging weighted prediction scores, or by combining learned features before the final classifier. A typical action-recognition pipeline is:

  • Sample RGB frames from a video for the spatial stream.
  • Compute optical-flow fields across consecutive frames for the temporal stream.
  • Run both streams and combine their evidence into labels such as “cycling,” “waving,” or “opening a door.”
Why it matters

Appearance alone cannot reliably recognize motion: a photo of a person holding a bat could depict baseball, cricket, or no action at all. Motion alone can lose useful context, especially when the camera moves or the scene is crowded. Two-stream networks combine these strengths for video search, sports analysis, surveillance, and autonomous-vehicle perception. Their main trade-off is that calculating optical flow can be expensive and imperfect; newer video transformers and 3D convolutional models frequently learn motion directly from frames. Still, the two-stream idea established a central lesson in video vision: recognizing actions requires reasoning about both content and change over time.

A Two-Stream Network is a video-recognition architecture that processes appearance and motion separately: one stream analyzes RGB frames, while the other analyzes temporal motion representations such as optical flow. Their predictions are fused to recognize actions or events. This separation lets models capture both what is visible and how it changes over time, improving video action recognition.

Imagine watching a football match with two helpers. One helper looks at each frame like a photograph: jerseys, players, the ball, and the field. The other pays attention to movement: who is running, where the ball is travelling, and whether someone is kicking or falling.

A Two-Stream Network gives an AI these two complementary views of a video. One stream focuses on what is visible in individual images, while the other focuses on changes and motion between them. Together, this helps the AI tell similar-looking moments apart—for example, a person standing beside a bicycle versus actually riding it. It matters because actions are defined not just by objects, but by how they move over time.