In Part 1, we saw the problem: uncompressed 1080p video runs at roughly 180MB per second. A 2-minute ad would be over 20GB. Nobody streams, stores, or emails files that size. So every video you’ve ever watched has been compressed — and understanding how explains a lot about why video pipelines behave the way they do.
The insight: video is full of redundancy
Compression works because video frames are rarely all that different from one another. Two kinds of redundancy get exploited:
Spatial redundancy — within a single frame, neighboring pixels are often similar (a blue sky, a plain wall). You don’t need to store every pixel independently; you can describe regions efficiently.
Temporal redundancy — between frames, most of the scene usually doesn’t change from one frame to the next. A talking head against a static background, for example, has a moving mouth but a background that’s identical frame after frame. Storing that background over and over is wasteful.
Modern video codecs (like those in the MPEG family — H.264/AVC, H.265/HEVC) are built to exploit both. This is where I-frames, P-frames, and B-frames come in.
The three frame types
I-frame (Intra-coded frame) A complete, standalone image — compressed the way a JPEG is, using only spatial redundancy. It doesn’t depend on any other frame. Think of it as a “reset point” or keyframe.
P-frame (Predicted frame) Instead of storing a full image, a P-frame stores only the difference from the previous frame (usually the most recent I-frame or P-frame). If the background hasn’t moved, that data barely takes any space at all — it just says “reuse what came before, adjust for this.”
B-frame (Bidirectional predicted frame) Goes a step further — it references both a previous and a future frame to predict its content. Because it can borrow information from either direction, B-frames are typically the most compressed of the three.
Why the order matters: the GOP
GOP
These frames are organized into a repeating pattern called a GOP — Group of Pictures. A typical GOP might look like:
I B B P B B P B B P …
- The I-frame anchors the group.
- P-frames periodically update what’s changed.
- B-frames fill in between, referencing both sides.
This is exactly why, if you’ve ever seen a video glitch and freeze into a blocky, smeared image during a bad network connection — that’s a lost or corrupted P/B-frame with nothing solid to reference, or a dropped I-frame breaking the whole chain until the next one arrives.
Why this matters practically
- Seeking/scrubbing in a video player is fast at I-frames and slower elsewhere — players often jump to the nearest I-frame first, then decode forward.
- Shorter GOPs (more frequent I-frames) mean better seek accuracy and more resilience to dropped frames, but larger file sizes.
- Longer GOPs compress more efficiently but recover more slowly from errors, and are less scrub-friendly.
- This is also why frequent cuts, fast motion, or heavy scene changes compress worse — every scene cut essentially forces something close to a full I-frame’s worth of new data, breaking the “not much has changed” assumption the whole system relies on.
By combining spatial compression (within I-frames) and temporal compression (across P and B frames), codecs like H.264 typically shrink raw video by 90–95% or more — turning that 180MB-per-second stream into something a phone can stream over LTE.
That compressed stream still needs to be packaged for delivery — synced with audio, subtitles, and metadata, and organized so a player can start playing before the whole file downloads. That’s where the MP4 container comes in, and it’s where Part 3 picks up.
This is Part 2 of a series on how video actually works under the hood — from raw bits to the compressed, streamable files we use every day.
Note: Drafted with help from Claude and Gemini — ideas, opinions, and the GPU bills are mine 😊