Data storage units, file sizes and compression
Questions on this topic are often calculations: how big a file is, or how many files fit in a space. The other half is explaining why files are compressed and which method suits which file.
The units
A bit is a single 0 or 1, a nibble is 4 bits and a byte is 8 bits. From there the syllabus uses binary prefixes, where each step is 1024 times the one before: 1 KiB = 1024 bytes, then MiB, GiB, TiB, PiB and EiB.
Calculating file sizes
An image file’s size in bits is its width in pixels × its height in pixels × its colour depth (bits per pixel). A sound file’s size in bits is its sample rate (samples per second) × its sample resolution (bits per sample) × its length in seconds.
Both give bits, so divide by 8 for bytes, then by 1024 for each step up to KiB, MiB and so on. For example, a 1024 × 768 image with a 16-bit colour depth is 1024 × 768 × 16 = 12 582 912 bits, which is 1 572 864 bytes, or 1536 KiB, or 1.5 MiB.
Why compress?
- Files take up less storage space.
- They are quicker to send and download, and use less bandwidth.
- Streaming works with less buffering.
Lossy and lossless compression
Lossy compression permanently removes data the user is unlikely to notice, such as detail the eye or ear can’t pick up, or by lowering the resolution or colour depth. The original can’t be rebuilt, but the file becomes much smaller. It suits photos, music and video.
Lossless compression makes the file smaller without losing any data, so the original is rebuilt exactly. It suits text, program files and spreadsheets, where one changed character would matter.
Run-length encoding
Run-length encoding (RLE) is a lossless method. It replaces a run of the same value with one copy of the value and a count. A row of pixels WWWWWWBBBW could be stored as 6W 3B 1W. It only saves space when there are long runs of repeated values; with few repeats the encoded version can be bigger than the original.
Where marks go
- Forgetting to divide by 8 to turn bits into bytes.
- Dividing by 1000 instead of 1024 when the question uses KiB, MiB or GiB.
- Describing lossy compression as “reducing quality” without saying data is permanently removed.
- Saying lossless compression “removes unnecessary data”; it removes no data at all.
- Not showing the working in a calculation, so no marks can be given for method.