Learn · Multimedia forensics

JPEG & PNG

The two most common image formats are both chains of self-describing blocks. Learn to walk the chain and you can tell where a picture really ends, which encoder wrote it, and whether a block was damaged.

1 Markers and segments

A JPEG is a sequence of segments. Each starts with FF and a marker byte. Apart from a few stand-alone markers (SOI, EOI, RST0–7), the marker is followed by a big-endian 16-bit length that counts itself and the payload, but not the two marker bytes:

next marker = marker offset + 2 + length FF D8 SOI start of image (no length) FF E0–EF APP0–APP15 application data: JFIF, Exif, XMP, ICC profile, … FF DB DQT quantisation tables FF C0 SOF0 frame header: precision, height, width, components (C2 = progressive) FF C4 DHT Huffman tables FF DA SOS start of scan, followed by the compressed data FF D9 EOI end of image (no length)

The sample photo below comes from a (fictional) camera. Its first segment after SOI is APP1 with the EXIF data, 4,738 bytes long. Every group ends with a → Next marker row: click it and check that it lands exactly on the next FF.

Try this: find the real image size in the SOF0 segment (height first, then width, both big-endian). The picture is 320 × 240 pixels. Then look at the three component entries: Y is sampled 2×2, Cb and Cr 1×1. That is 4:2:0 subsampling, the usual camera setting: colour is stored at half resolution in both directions.

2 The scan and the end of the image

After the SOS header comes the entropy-coded data: the compressed pixels, Huffman-coded bit by bit. It has no length field. A decoder reads until it meets a marker, normally EOI. So that an FF data byte can never look like a marker, the encoder writes it as FF 00 (byte stuffing). In the sample the scan is 11,000 bytes long and contains 193 stuffed pairs.

FF 00 data byte FF (stuffing) FF D0 … FF D7 restart markers, may appear inside the scan FF D9 end of image FF + any other byte a new segment (progressive JPEGs have several scans)

The first FF D9 after the last scan is the end of the picture. A viewer stops there. Whatever follows is invisible, but it is still in the file and still counted in the file size: the hex view flags it as trailing data. The Tampering page shows a JPEG with an archive appended after EOI.

Header and footer. FF D8 FF at the start and FF D9 at the end are the signatures carving tools search for in raw data. Because the scan has no length, a carver cannot know where a JPEG ends without either decoding it or searching for FF D9.

3 Quantisation tables

JPEG cuts the image into 8×8 blocks, transforms each block into 64 frequency coefficients (DCT), and divides every coefficient by the matching entry of a quantisation table before rounding. Large divisors throw more detail away. This rounding is where JPEG loses information, and the "quality" slider of a program only chooses the table.

The table is stored in zigzag order: from the top-left corner (the average brightness, "DC") diagonally towards the highest frequencies at the bottom right. The view rearranges it into the 8×8 grid.

IJG scaling (libjpeg, Pillow, GIMP, …): scale = quality < 50 ? 5000 / quality : 200 − 2 × quality entry = clamp(1, 255, ⌊(base entry × scale + 50) / 100⌋)

Most software does not invent tables: it scales the two example tables from the JPEG standard with the formula above. That makes it possible to estimate a "quality" from any table, and to tell whether a table is a scaled standard table at all. The sample camera's luminance table starts 4 5 5 7 7 7 8 9 in zigzag order; the closest IJG quality is about 87, but it does not match any IJG table exactly. That is typical for camera firmware. The Attribution page uses this as a fingerprint.

4 PNG chunks

PNG is lossless and has a simpler, fully length-prefixed structure: an 8-byte signature, then chunks. Each chunk is length (4, big-endian) · type (4 ASCII letters) · data · CRC-32 (4). The CRC covers type and data, so a single changed byte is detectable.

The case of the first letter of the type says whether a decoder must understand the chunk: IHDR, IDAT, IEND are critical; tEXt, pHYs, eXIf are ancillary and may be ignored. The two tEXt chunks here hold a Title and a Comment as plain "keyword NUL text".

Try this: the IHDR data is 13 bytes: width (200 px), height (150 px), bit depth, colour type, compression, filter and interlace method. The CRC after it is computed over the 4 type bytes + 13 data bytes = 17 bytes. Like EOI in a JPEG, IEND ends the picture, and anything after it would be flagged as trailing data.

Chunk listing (pngcheck-style)
No camera metadata in PNG? PNG had no standard place for EXIF until the eXIf chunk (2017), and many programs still drop it. Screenshots and edited images are often PNG; a camera photo saved as PNG has usually lost its EXIF on the way.

5 Try it yourself

Answer from the hex views above. Numbers can be typed in decimal or hex (0x…).

The APP1 segment of the camera photo starts at file offset 0x2. At which offset does the next segment start?

Read the big-endian length after FF E1. It counts itself and the payload, but not the two marker bytes.

The length at 0x4 is 12 82. Big-endian: 0x1282 = 4,738.

Next marker = 0x2 + 2 + 0x1282 = 0x1286 (4,742). There the file continues with FF DB, the first DQT segment.

How many bytes of the file does one DQT segment occupy, marker included?

The DQT at 0x1286 holds one 8-bit table: 1 byte precision/id + 64 values. Add the length field and the marker.

The length after FF DB is 00 43 = 67: 2 bytes for the length itself + 1 byte precision/id + 64 table entries.

With the two marker bytes: 2 + 67 = 69 bytes. Check: 0x1286 + 69 = 0x12CB, where the second DQT starts.

Which quantisation table does the Cr component of the camera photo use?

Each component entry in SOF0 is 3 bytes: component id, sampling factors (high nibble horizontal, low nibble vertical), table id.

The third component entry at 0x1320 is 03 11 01: id 3 (Cr), sampling 1×1, quantisation table 1.

Cb (02 11 01) shares the same table. Only Y (01 22 00) uses table 0. The file has exactly two DQT segments, ids 0 and 1.

How many bytes of compressed pixel data does the IDAT chunk of the PNG sample hold?

The chunk length is the 4 bytes before the type IDAT, big-endian, and counts only the data (not type or CRC).

The IDAT chunk starts at 0x9C with 00 00 0A 47. Big-endian: 0x0A47 = 2,631 bytes.

The whole chunk is 4 + 4 + 2,631 + 4 = 2,643 bytes, so the next chunk (IEND) starts at 0x9C + 2,643 = 0xAEF.

These are the first 16 bytes of the camera photo: SOI, the APP1 marker and length, Exif\0\0 and the start of the TIFF header. Edit them and watch the segment walk below:

  • Change the length 12 82 to 12 83. Where does the walk land now, and what does the parser make of the rest of the file?
  • Reset, then change the marker FF E1 to FF E0 (APP0). The bytes are still "Exif", but the parser no longer reads them as a TIFF structure: tools decide by marker and identifier.
  • Reset, then set the length to 00 10. The walk jumps into the middle of the EXIF data. Why is the result "unexpected data" rather than a wrong segment?