Part 1 — Photographs: reading the trap before you hear it
A close reading of how TOEIC Part 1 builds its four options — tense and voice traps, sound-alikes, and the state-versus-action distinction that decides most of the six questions.
Scan in a fixed order, every time
You get a few seconds of silence before the narrator starts, and nothing is printed on screen except the photo — no text version of the four statements, so everything you hear has to be matched against what you already noticed. Run the same scan every time: main subject first (a person if there is one, otherwise the dominant object), then what that subject is actually doing, then the secondary objects and the setting. If no one in the photo is doing anything specific, that is itself information — it tells you the correct statement is likely to describe the setting or a state, not an action.
The tense-and-voice trap, in detail
- Active present continuous ('is operating', 'is pressing') describes someone doing something right now, in the frame — this is the most common correct-answer form when a single person is the clear subject of the photo.
- Passive present continuous ('is being wiped down', 'is being loaded') needs someone visibly performing that exact action at this instant. A plausible object in the scene is not enough — if no hands are on it, the option is wrong no matter how natural it sounds.
- Present perfect passive ('has been placed', 'have been arranged', 'has been erected') describes a finished result, with no one necessarily doing anything now. This is usually the safe, correct form for photos where nothing is actively happening, or where the action can't be pinned to a specific person — distant shots, group shots, stacked or arranged objects.
- Stative verbs (sit, stand, wear, hold, face, overlook) describe a condition, not an action in progress. Don't force these into an 'is X-ing' reading in your head, and don't expect to find an active agent for them — 'a picture is hanging on the wall' needs no one hanging it right now.
Sound-alike distractors are deliberate, not incidental
- The test writers build at least one option per photo around a word that sounds like the correct one but changes the meaning entirely: work/walk, copy/coffee, fund/found, seat/sit, lead/lead (verb/noun stress), pass/past.
- Homophone pairs are especially common on verbs: writing/righting is a real example from this site's own bank (a woman writing in a notebook, distractor claims she's 'righting an overturned chair').
- Near-rhymes on technical vocabulary show up too — tamping/stamping (a barista pressing a tamper into a portafilter, distractor claims she's 'stamping a customer's loyalty card'). These pass at conversational speed if you're listening passively; you have to be listening for the specific action, not just recognising a familiar-sounding syllable.
- Because the four statements are read aloud back to back with no repeat, a sound-alike only needs to occupy your ear for half a second to do its job. Treat any option that 'sounds right' with extra suspicion, not less.
Four ways a distractor gets built
- Right object, wrong action: the item is genuinely in the photo, but the verb attached to it is false. Crates stacked against a wall become 'he's loading crates onto a delivery truck' — the crates are real, the loading isn't happening.
- Right action, wrong object: the true action is close by, but the option attaches it to the wrong thing — 'the barista is pouring milk into a jug' when she's actually working the tamper, not the milk jug.
- Plausible but absent action: an action that would fit the general scene but isn't shown at all — 'the passengers are boarding the plane' for people sitting in a waiting area, or 'a crane is lifting a steel beam' for a construction site where no crane is even visible.
- True detail, wrong statement type: the option correctly names something visible but phrases it as an action when it's a state, or vice versa — describing scaffolding as something someone is currently building rather than as 'has been erected', a completed condition.
Group shots and distant shots default to state descriptions
When a photo is taken from far away, or shows several people at once, you usually can't tell exactly what any one individual is doing with their hands — and the test knows this. In these cases the correct answer is very often a passive present-perfect statement about the overall scene ('scaffolding has been erected around the building') rather than a claim about a specific person's action, because a specific-action claim can't be verified at that distance or in that crowd. When you sense you're looking at a wide or group shot, mentally raise the odds that the answer is a state description before the audio even starts.
One practical habit that pays off immediately
As each statement plays, don't wait for the fourth option before deciding — mark A, B, C, or D as a tentative keep or reject the instant you hear it, based on the scan you already did. By the time D finishes, you're choosing between at most two live candidates instead of replaying all four from memory. On a six-question part with no time to relisten, that habit alone recovers points that pure vocabulary study won't.
