The big idea
Everything on a computer is stored as bits. Compression is the art of saying the same thing with fewer of them.
By the end of this part you can
- explain what compression is and why we need it
- convert between bits, bytes, kilobytes and megabytes
- calculate a compression ratio and a percentage saving
Why bother making files smaller?
A single photo from a modern phone camera has about 12 million pixels. Each pixel needs 3 bytes (one each for red, green and blue). That's 36 million bytes (36 MB) for one photo, before compression. If every photo stayed that size, a 128 GB phone would fill up after roughly 3,500 photos, and sending one to a friend on a slow connection would take minutes.
In practice that same photo saved as a JPEG is usually 3 to 5 MB. Compression is why you can stream a movie, send voice messages and keep thousands of songs on your phone.
Compression saves four things:
- Storage
- More files fit on the same drive, phone or server.
- Bandwidth
- Less data to send means faster downloads and smoother streaming.
- Money
- Companies pay for storage and data transfer. Mobile plans have data caps.
- Energy
- Moving and storing fewer bits uses less electricity in data centres.
A quick refresher on units
A bit is a single 0 or 1. A byte is 8 bits. From there we go up in thousands:
| Unit | Size | Roughly enough for |
|---|---|---|
| 1 byte (B) | 8 bits | one letter of plain text |
| 1 kilobyte (KB) | 1,000 bytes | a short paragraph |
| 1 megabyte (MB) | 1,000 KB | a one-minute MP3 |
| 1 gigabyte (GB) | 1,000 MB | about an hour of HD streaming |
| 1 terabyte (TB) | 1,000 GB | a large external hard drive |
Speeds are measured in bits per second, not bytes. A 50 Mbps internet connection moves 50 million bits each second, which is 6.25 million bytes. Divide by 8 to go from bits to bytes.
b means bits and is used for speeds: 50 Mbps is 50 megabits per second. A capital B means bytes and is used for storage and file sizes: 50 MB is 50 megabytes. So a 50 MB file downloaded over a 50 Mbps connection takes about 8 seconds, not 1. Advertisers rely on people missing that difference.Encoding and decoding
Compression always has two halves. An encoder squeezes the data down. A decoder turns it back into something you can see or hear. Together they're called a codec (coder-decoder). MP3, JPEG and H.264 are all codecs. Both ends must agree on the rules, which is why an old phone sometimes can't play a video made with a newer codec.
Measuring how well it worked
The compression ratio compares the original size to the compressed size:
compression ratio = original size ÷ compressed size
A 10 MB file compressed to 2 MB has a ratio of 10 ÷ 2 = 5:1 ("five to one"). You can also express it as a percentage saving:
saving = (1 − compressed ÷ original) × 100%
For the same file: (1 − 2 ÷ 10) × 100% = 80% smaller. Make sure both sizes are in the same unit before you divide.
Enter an original and compressed size. Try the presets to see how different real files compare.
Dark bar: original. Orange bar: compressed, drawn to scale.
Watch
Lossless or lossy
Every compression method makes one big choice: keep everything, or throw some of it away.
By the end of this part you can
- explain the difference between lossless and lossy compression
- explain redundancy and irrelevance
- choose the right type of compression for a situation, and justify it
Lossless: nothing is lost
Lossless compression shrinks a file so that when it's decoded you get back exactly the original, bit for bit. It works by finding redundancy: patterns and repetition that can be written more efficiently.
Think of the sentence "the cat sat on the mat and the cat sat on the rug". Instead of writing "the cat sat on the" twice, you could write it once and say "repeat that here". Nothing is lost; you've just found a shorter way to describe it.
Lossless is essential when every bit matters: text documents, program code, spreadsheets, databases, and medical scans. Imagine a bank statement where "compression" changed $1,000 to $1,001. Formats: ZIP, PNG, GIF, FLAC, ALAC.
Lossy: throw away what you won't notice
Lossy compression permanently removes some data. It targets irrelevance: detail that humans can't see or hear, or won't notice. Your eyes are much better at seeing changes in brightness than changes in colour; your ears can't hear a quiet sound right next to a loud one. Lossy codecs exploit these limits.
Because it can discard data, lossy compression gets far smaller files, often 10 to 100 times smaller. The trade-off: once data is thrown away it can never be recovered. Formats: JPEG, MP3, AAC, H.264, WebP (which can do both).
| Lossless | Lossy | |
|---|---|---|
| Decoded file | Identical to the original | An approximation of the original |
| Removes | Redundancy (repetition, patterns) | Redundancy and irrelevance (detail we won't notice) |
| Typical ratio | About 2:1 to 3:1 for most files | 10:1 to 100:1 or more |
| Quality control | None needed, always perfect | A quality setting or bitrate chooses the size/quality trade-off |
| Use for | Text, code, data, logos, screenshots, archiving | Photos, music, streaming video, voice calls |
So which is better?
Neither. It depends on what the data is and what it's for. Ask yourself:
- Would changing a single bit cause a problem? If yes, lossless.
- Is a human going to look at it or listen to it? If yes, lossy is usually fine.
- Will it be edited again later? Keep a lossless master.
- Is bandwidth or storage very limited? Lossy, with the quality set as high as the budget allows.
Decide which type of compression suits each situation. Read the feedback, especially if you disagree.
Think deeper: is a "lossless" photo really perfect?
A camera sensor already loses information: it samples light at a fixed number of pixels and a fixed number of brightness levels. So a "lossless" PNG of a photo is a perfect copy of the digital photo, not of the real scene. Converting anything from the real world into bits (digitising) always involves choices about how much detail to keep. You'll see this again with audio sampling in Part 4.
Lossless techniques
Three classic ways to find redundancy. Real formats like ZIP and PNG combine them.
By the end of this part you can
- encode and decode data using run-length encoding (RLE)
- explain how dictionary compression replaces repeated patterns
- build a Huffman code and calculate how many bits it saves
Technique 1: Run-length encoding
A run is a sequence of the same value repeated. RLE replaces each run with a count and the value. This row of pixels (W = white, B = black):
W W W W W W B B B W W W W
becomes 6W 3B 4W. Thirteen values have become three pairs.
RLE is brilliant for images with big areas of flat colour: icons, simple logos, fax pages, old-school game sprites. It's terrible for photos or noisy data, where neighbouring values are rarely identical. Then every "run" is a length of 1, and storing a count as well as the value makes the file bigger.
Paint an 8×8 image. Each pixel is stored as 1 byte (8 bits). In the RLE version each run costs 2 bytes: one for the count, one for the colour. Runs restart at the end of each row.
Challenge: find a picture where RLE is exactly the same size as the original. Then explain why the checkerboard is the worst case.
Technique 2: Dictionary compression
Instead of runs of one value, dictionary methods look for repeated sequences: whole words, phrases or byte patterns. Each repeat is swapped for a short code, and a dictionary records what each code means.
The famous LZ family of algorithms (named after Abraham Lempel and Jacob Ziv, 1977) does this cleverly: instead of a separate dictionary, it points back to text it has already seen, like "copy the 12 characters that appeared 40 characters ago". ZIP files and PNG images both use a method called DEFLATE, which combines an LZ-style dictionary with Huffman coding.
Words of 3 or more letters that repeat are replaced by a token like #1. A word only goes in the dictionary if it actually saves space, since the dictionary has to be stored too. Edit the text and see what happens.
Dictionary
Compressed text
Technique 3: Huffman coding
Normal text uses a fixed-length code: every character takes 8 bits, whether it's a common "e" or a rare "z". David Huffman, a student at MIT in 1952, realised it's smarter to use a variable-length code: give common characters short codes and rare characters long ones. Morse code uses the same idea (E is a single dot).
How to build a Huffman code:
- Count how often each character appears (its frequency).
- Treat each character as a little tree with one node. Take the two with the lowest frequencies and join them under a new parent whose frequency is their total.
- Repeat until everything is joined into one tree.
- Label every left branch 0 and every right branch 1. A character's code is the path from the top of the tree down to it.
Because every character sits at the end of a branch, no code is the start of another code. This is called the prefix property, and it means the decoder never gets confused about where one character ends and the next begins, even with no gaps between codes.
Type a short message (up to about 60 characters works best for the tree). The table shows each character's frequency and code. Spaces are shown as ␣.
The encoded message
Worked example: "BANANA" by hand
Frequencies: A = 3, N = 2, B = 1.
Join the two smallest: B (1) + N (2) make a node worth 3. Now join that node (3) with A (3) to make the root (6).
Labelling left 0, right 1 gives one possible result: A = 0, B = 10, N = 11. (If you put A on the right instead you'd get different codes of the same lengths, which is equally valid.)
BANANA → 10 0 11 0 11 0 = 9 bits. As 8-bit text it would be 6 × 8 = 48 bits. Saving: (1 − 9 ÷ 48) × 100% ≈ 81%.
The catch: the decoder needs the code table too, so a real file stores the tree as well. For a six-letter word the tree costs more than it saves, but for a whole book it's tiny by comparison.
Watch
Audio
How a sound wave becomes numbers, why raw audio is huge, and how MP3 uses the quirks of your ears to shrink it.
By the end of this part you can
- explain sample rate and bit depth
- calculate the size of an uncompressed audio file
- explain how lossy audio codecs use psychoacoustics, including masking
From wave to numbers
Sound is a wave of air pressure. A microphone turns it into a smoothly changing voltage. To store it digitally, an analogue-to-digital converter measures that voltage thousands of times per second. Each measurement is a sample.
Two settings decide the quality and the size:
- Sample rate
- How many samples per second, in hertz (Hz). CD audio uses 44,100 Hz (44.1 kHz).
- Bit depth
- How many bits store each sample. 16 bits gives 65,536 possible levels. More bits means less rounding error.
- Channels
- Mono is 1 channel, stereo is 2 (left and right).
Why 44.1 kHz? Humans hear roughly 20 Hz to 20,000 Hz. The Nyquist theorem says you need to sample at at least twice the highest frequency you want to capture. Twice 20 kHz is 40 kHz, plus a little safety margin.
If you sample too slowly, there aren't enough measurements to describe a fast wave, and the samples end up matching a slower wave just as well. When the computer plays them back it produces that slower wave instead, so a high note is recorded as a lower note that was never there. This error is called aliasing, and you can hear it for yourself in the next activity.
The grey curve is the real sound wave. Orange dots are samples, snapped to the nearest level the bit depth allows. The dark line is what the computer can rebuild. Then press Play to hear your settings.
Hear it
Keep the volume low. Try a 2,500 Hz tone at 3,000 Hz sample rate: you'll hear a much lower note. That's aliasing, because 3,000 Hz is less than twice 2,500 Hz. Then drop the bit depth to 2 and listen to the buzz of rounding errors.
How big is raw audio?
Uncompressed audio (a WAV file, for example) is easy to calculate:
size in bits = sample rate × bit depth × channels × seconds
Divide by 8 for bytes. A 3-minute song at CD quality: 44,100 × 16 × 2 × 180 = 254,016,000 bits = 31,752,000 bytes ≈ 31.75 MB. That's why nobody streams raw CD audio.
The number of bits per second is the bitrate. For CD audio it's 44,100 × 16 × 2 = 1,411,200 bits per second, or about 1,411 kbps.
Same length, compressed
Lossless audio: FLAC
FLAC (Free Lossless Audio Codec) predicts each sample from the ones before it, then stores only the small difference between the prediction and the real value. Small numbers need fewer bits. Music typically shrinks to around 50 to 60% of its WAV size, with zero loss. Musicians and audiophiles use it for archiving.
Lossy audio: MP3, AAC and Opus
Lossy audio codecs rely on psychoacoustics: the science of how we actually perceive sound. They split the sound into many frequency bands, then use a psychoacoustic model to decide which parts you won't notice missing. Main tricks:
- Cutting what's out of range. Very high frequencies (above about 16 to 18 kHz) are hard for most people to hear, especially as they get older. At low bitrates they're simply removed.
- Frequency masking. A loud sound hides quieter sounds at nearby frequencies. If a cymbal crashes, the codec doesn't waste bits on a soft sound right next to it in pitch.
- Temporal masking. A loud sound also hides quiet sounds that happen a few milliseconds before or after it.
- Joint stereo. Left and right channels are often nearly identical, so the codec stores what they share once, plus the small differences.
- Huffman coding packs whatever's left, losslessly. Your Part 3 skills show up again.
The quality setting is the bitrate. 128 kbps MP3 is about 11 times smaller than CD audio; 320 kbps sounds essentially identical to the original for most listeners. Newer codecs like AAC (used by Apple) and Opus (used by many voice and video apps) get better quality than MP3 at the same bitrate.
A loud 1,000 Hz tone plays with a quiet second tone. Can you hear the quiet tone? Compare a quiet tone that's close in pitch with one that's far away. Volume low, headphones help.
Most people find the first two sound almost the same, but the far tone is clearly audible. An MP3 encoder would throw away the close quiet tone and keep the far one. That's irrelevance in action.
Watch
Images
Pixels, colour depth, and the six-step recipe that makes JPEG so effective on photos.
By the end of this part you can
- calculate the size of an uncompressed bitmap image
- choose between PNG and JPEG for a given image, and explain why
- describe the steps of JPEG compression and identify the lossy steps
Images are grids of numbers
A bitmap image is a grid of pixels. The resolution is its width × height in pixels. The colour depth is how many bits store each pixel's colour:
| Colour depth | Possible colours | Used for |
|---|---|---|
| 1 bit | 2 (21) | black and white, like a fax |
| 8 bits | 256 (28) | GIFs and simple graphics |
| 24 bits | 16,777,216 (224) | photos: 8 bits each for red, green and blue ("true colour") |
size in bits = width × height × colour depth
A 1920 × 1080 screenshot in 24-bit colour: 1920 × 1080 × 24 = 49,766,400 bits ÷ 8 = 6,220,800 bytes ≈ 6.2 MB.
Lossless images: PNG and GIF
PNG first applies a "filter" that stores each pixel as its difference from its neighbours (in smooth areas those differences are small and repetitive), then compresses the result with DEFLATE, the same LZ + Huffman combination used in ZIP. GIF limits the image to a palette of 256 colours and uses a dictionary method called LZW.
Lossless formats are perfect for logos, diagrams, icons, screenshots and text: sharp edges and flat colour, where lossy artefacts would be obvious. They are poor for photos, where no two neighbouring pixels are quite the same.
Lossy images: how JPEG works
JPEG is designed for photographs. It uses two facts about human vision: we notice brightness detail far more than colour detail, and we notice gradual changes more than tiny, fine detail. Here's the recipe:
- Change colour space. Convert RGB into Y (brightness, called luma) plus Cb and Cr (two colour channels, called chroma). Lossless on its own.
- Chroma subsampling. Store colour at a lower resolution, often a quarter of the pixels, while keeping full brightness detail. Lossy.
- Split into 8×8 blocks. Each block of 64 pixels is processed separately.
- Discrete cosine transform (DCT). Rewrite each block as a mix of 64 standard patterns, from smooth (top-left) to very fine stripes (bottom-right). Lossless on its own: it's just a different way to describe the same block.
- Quantisation. Divide each pattern's amount by a number and round it. The fine-detail patterns are divided heavily and often round to zero. This is the main lossy step, and the "quality" slider controls how harsh it is. Lossy.
- Pack it up. Read the values in a zigzag so the zeros bunch together, then use RLE and Huffman coding. Lossless. Parts 3 and 5 meet!
Your browser really encodes this image as a JPEG every time you move the slider. The sizes are genuine. Turn on zoom to see the 8×8 blocks appear at low quality.
Look for blocking (visible squares), ringing (ripples next to sharp edges, see the lettering) and colour bleeding across edges. Where does quality start to look bad to you, and how big is the file at that point?
Step 2 of JPEG in action. The middle image keeps brightness at full resolution but stores colour at a much lower resolution. Can you spot the difference from the original?
The third image throws away the same amount of brightness data instead. Much worse, right? That's why JPEG shrinks colour, not brightness.
Steps 4 and 5 up close. The DCT turns 64 pixels into 64 pattern amounts (right grid). Keep only the first few in zigzag order and rebuild the block. How few can you keep before it looks wrong?
Notice that the smooth gradient looks fine with just a handful of values, but the fine checkerboard needs nearly all 64. Photos are mostly smooth, which is why JPEG works so well on them, and so badly on crisp text.
Watch
Video
Video is a stack of images plus sound. It would be impossibly big without compression, so it uses every trick so far and adds one more: only store what changes.
By the end of this part you can
- calculate the size and bitrate of uncompressed video
- explain the difference between spatial and temporal compression
- explain I-frames, P-frames and B-frames, and why fast motion hurts quality
Why video needs compression the most
Video is a sequence of still images called frames, shown quickly enough to look like motion. The frame rate is frames per second (fps): 24 fps for films, 30 or 60 fps for TV, games and phone video.
size in bits = width × height × colour depth × frame rate × seconds
One second of 1080p at 30 fps in 24-bit colour: 1920 × 1080 × 24 × 30 = 1,492,992,000 bits, about 1.5 gigabits every second. A typical home connection can't come close to that, yet streaming services deliver good-looking 1080p at around 5 Mbps. That's a ratio of roughly 300:1.
Two kinds of video compression
Spatial compression (also called intra-frame) squeezes each frame on its own, much like a JPEG. It removes redundancy within one picture.
Temporal compression (inter-frame) removes redundancy between frames. In most video, very little changes from one frame to the next: a person talking in front of a still background might only move their mouth and eyes. So why store the whole background 30 times per second?
- I-frame (keyframe)
- A complete picture, compressed like a JPEG. Needed at the start, at scene cuts, and every few seconds so you can skip around.
- P-frame (predicted)
- Stores only what changed since the previous frame, plus motion vectors: "this block moved 4 pixels left".
- B-frame (bi-directional)
- Predicted from frames both before and after it. The smallest type, but the encoder has to look ahead.
- Group of pictures (GOP)
- One I-frame and the P- and B-frames that depend on it, until the next I-frame.
Streaming video usually uses a bitrate budget: a limit on how many bits per second can be sent. When a scene has lots of unpredictable change, like snow, confetti, water spray or a crowd jumping, the P-frames suddenly need far more data. If the budget can't grow, the encoder has to throw away more detail, and the picture turns into a blocky smear. That's the reason behind the Tom Scott video below.
Each frame is 48 × 27 blocks. I-frames (yellow bars) store every block. P-frames (teal bars) store only blocks that changed. If a frame needs more than the budget, it turns orange and the picture degrades. Turn on confetti.
Why does a camera pan cause trouble in this simple simulator, even though real codecs handle it quite well? (Hint: think about motion vectors.)
Codecs and containers
A video file like holiday.mp4 is a container: a box holding a compressed video stream, a compressed audio stream, subtitles and information to keep them in sync. MP4, MOV, MKV and WebM are containers. The codec is how each stream inside was compressed.
| Video codec | Where you'll meet it | Notes |
|---|---|---|
| MPEG-2 | DVDs, older digital TV | The 1990s standard. |
| H.264 (AVC) | Almost everywhere: web video, phones, Blu-ray | Very widely supported. |
| H.265 (HEVC) | 4K streaming, recent phones | Roughly the same quality as H.264 at about half the bitrate, but licensing is complicated. |
| VP9 and AV1 | YouTube, Netflix and other streaming | Royalty-free. AV1 is newer and more efficient but takes more computing power to encode. |
Streaming services also use adaptive bitrate streaming. Each video is encoded several times at different resolutions and bitrates, chopped into short chunks. Your player measures your connection and switches between versions chunk by chunk. That's why a video sometimes goes blurry for a few seconds and then sharpens again.
Watch
Squeeze real files
Predict, test, record and explain. You'll compress real files with the same methods your computer uses every day.
This practical has three stations. Every result you get is saved on this page and goes into your practical report at the bottom. Your files never leave your computer: all the compression happens inside your browser, and nothing is uploaded anywhere.
No files of your own? Every station has sample files you can use instead. Using some of your own as well will give you more interesting results.
Success criteria
- I made predictions before testing, and compared them with my results
- I tested at least five different types of file with lossless compression and proved the output was identical
- I measured how file size and quality change with lossy image and audio compression
- I explained my results using the terms redundancy, irrelevance, lossless, lossy and compression ratio
Your details
Station 1: Lossless compression with gzip
gzip uses DEFLATE, the same LZ-dictionary-plus-Huffman method inside ZIP files and PNG images. It works on any file. Your job: find out which kinds of file it can shrink, and prove that nothing is lost.
Before testing, predict how much gzip will shrink each type of file. Don't change these after you've seen the results; your report compares them.
Add the sample files, your own files, or both. Each file is compressed, then decompressed and compared byte by byte with the original.
What does "identical" mean here?
After compressing, the page decompresses the file again and checks every single byte against the original. It also shows a SHA-256 fingerprint of each: a 64-character code calculated from the file's contents. Change even one bit of a file and its fingerprint changes completely. If the fingerprints match, the files are identical. That's your proof that gzip is lossless.
Station 2: Lossy image compression
Now you'll save one photo at several JPEG quality settings (and as PNG and WebP) and measure exactly what's lost. The difference map shows where the compressed image differs from the original: black means identical, bright means a big difference (made 8 times brighter so you can see it).
What happens if an image is saved as a JPEG over and over? Row 1 simply re-saves it. Row 2 nudges the image by one pixel before each save, like a small edit or crop, so the 8×8 blocks never line up the same way twice.
Station 3: Lossy audio compression
First, find out how compressed an audio file you already have is by comparing it with the raw PCM data inside it. Then re-encode 8 seconds of it at three bitrates and listen for the difference. Headphones help.
Encoding happens in real time, so it takes 8 seconds. Your browser chooses the codec (usually Opus or AAC) and may not hit each bitrate exactly; the table shows what it actually produced.
Analysis and conclusions
Refer to your own measurements in every answer. Numbers make an explanation convincing.
Practical report
Compression consultant
Use everything from Parts 1 to 6 to advise a real-world client with a slow internet connection.
This task has four parts. Parts A, B and C are marked automatically when you lock them in, and you can only lock them in once, so check carefully. Part D is written and marked by your teacher against the rubric at the bottom. When you're finished, download your PDF and submit it to your teacher.
Your answers are saved in this browser as you go, but save your PDF before closing the page on a shared computer.
Your details
The client: Riverbend Community Arts Centre
Riverbend is a volunteer-run arts centre in a small town in north-west Tasmania. Their satellite internet gives them 25 Mbps download, only 1 Mbps upload, and a 50 GB monthly data allowance for uploads and downloads combined. They want to share five things online:
- The weekly radio show: a 45-minute spoken-word program with some background music, released as a podcast.
- Exhibition gallery: 40 photographs of paintings from the annual art prize, shown on their website.
- Logo and festival map: the centre's flat-colour logo, and a map of stall locations with small text labels.
- Festival promo video: a 3-minute 1080p video. It ends with a 20-second finale where confetti cannons fire over the crowd.
- Oral history archive: recordings of local elders telling the town's history, which the state library wants to preserve for at least 50 years.
Use 1 KB = 1,000 bytes, 1 MB = 1,000,000 bytes and 1 GB = 1,000,000,000 bytes throughout.
Part A: Knowledge check
10 marks. Choose the best answer.
Part B: Crunch the numbers
7 marks. Enter numbers only (no units). Round to 2 decimal places where needed. Show your working on paper or in Part D if your teacher asks.
Part C: Apply the techniques
4 marks. Use the colour letters W, N and Y exactly as in the RLE painter, with no spaces.
Part D: Your recommendations
Teacher-marked against the rubric. Write in full sentences and use correct terminology.
For each of Riverbend's five items, choose a file format and explain why it suits that item. Refer to lossless or lossy compression, what kind of data it is, and what it will be used for. (2–4 sentences each.)
Choose either MP3 audio compression or JPEG image compression. Explain step by step how it reduces file size, and explain how it takes advantage of the limits of human hearing or vision. (150–250 words)
The promo video will be streamed at a fixed bitrate of 8 Mbps. Explain what is likely to happen to the picture quality during the confetti finale, and why, using the terms I-frame, P-frame and bitrate. Suggest one practical way Riverbend could reduce the problem. (100–150 words)
Using your Part B answers, estimate whether Riverbend can upload four weekly podcasts (as 64 kbps MP3), the promo video (at 8 Mbps) and the 40 gallery photos (as 3.6 MB JPEGs) within their 50 GB monthly allowance, and whether uploading at 1 Mbps is realistic for a volunteer. Then discuss one wider impact of compression, such as access for rural communities, environmental cost, or cultural preservation. (150–250 words)
Rubric and self-assessment
Read each row, then choose the level you think your work shows. Your teacher makes the final judgement.