Twenty seconds to look at a photograph, then up to ninety seconds to describe it out loud. There is one recording and no retake, which makes the first ten seconds the most valuable ten seconds to rehearse on the whole test.
| Time | 90 seconds |
| Preparation time | 20 seconds |
| How many per test | 1 |
| Subscore | Speaking |
| Adaptive | No |
| Scored as | Speaking criteria: content, discourse coherence, fluency, grammar, lexis, pronunciation |
A photograph with a twenty-second preparation timer. When it ends, a Record Now button appears; the recording starts when you click it and runs for up to ninety seconds. The photo stays on screen throughout.
Ninety seconds is a long time to talk about one image. Most test takers run out of things to say at about thirty and then lose fluency marks filling the rest. The fix is a plan, made in the twenty seconds.
Do not rehearse sentences. Pick four things to talk about, in order, and let the sentences happen live. This is what stops the answer stalling at thirty seconds.
"This looks like a bus stop, and it has been taken from behind the two people in the foreground. On the left there is an older man in a dark coat carrying a bag in each hand, and on the right a woman in a patterned coat standing beside a wheeled suitcase. Both of them are facing away from the camera and looking out towards the road, so I would guess they are waiting for something rather than about to leave. Further along the kerb there is a cream and green bus with a few other passengers near it. The trees behind have no leaves, which makes me think it is winter, and nobody seems to be in a hurry, so the bus is probably not due for a few minutes yet."
Filler that says nothing — "so, um, in this picture we can see, um" — costs Fluency directly. Filler that carries content, like a location phrase, does not.