Skip to content
All study tips
Listening & Reading 6 min read

Part 2 — Question–Response: why the first three words decide everything

A deep dive into TOEIC Part 2's 25 question–response items — the question-type taxonomy, the traps built into every wrong option, and why the best answers often don't sound like answers at all.

No picture, no text, just three voices

Part 2 strips away every non-audio crutch the test gives you elsewhere. The screen shows a single generic instruction — choose the best response to the question or statement you hear — and nothing else: no photo to scan, no question stem to read, no answer choices printed for your eyes. The narrator reads one question or statement, then three spoken responses labeled A, B, and C, and you have only a few seconds after option C to mark an answer before the next item starts. Twenty-five items run back to back with no repeats and no way to slow the audio down, which makes Part 2 the purest listening-only stretch of the entire test.

The question-type taxonomy

  • Wh-questions (Who, What, When, Where, Why, How, How many, How long) demand content that fills that exact slot. 'When is the quarterly report due?' only accepts a time ('By the end of the week') — a location or a yes/no reply is automatically wrong, no matter how fluent it sounds.
  • Yes/no questions, led by an auxiliary (Did, Is, Do, Can, Has, Was), expect confirmation or denial, often with a supporting detail attached — 'Has the recall notice been sent to all affected owners?' answered by 'Not yet, we're still finalizing the mailing list' is a full, correct answer even though it never says the word 'no'.
  • Tag questions ('...isn't it?', '...doesn't it?') function exactly like yes/no questions despite the extra clause riding along — 'The patient's insurance is still valid, isn't it?' is answered with a direct confirmation, not with commentary on the embedded statement.
  • Negative questions ('Isn't the new assembly line supposed to start today?') trip up test-takers who assume the negative flips the logic. It doesn't — you still answer as if the question were positive: yes if it's true, no if it isn't.
  • Choice (alternative) questions built around 'or' reject a plain yes/no outright. 'Should we hold the training session in the main hall or the small conference room?' needs one of the two named options picked, ideally with a reason attached.
  • Embedded questions bury the real question inside a frame like 'Do you know...' or 'Could you tell me...'. 'Do you know when the contract needs to be signed?' is really asking 'when' — answer the embedded clause, not the frame verb, or you'll fall for a response about being personally acquainted with something.
  • Statements with no question word at all still need a reactive response: agreement, disagreement, a next step, or an appropriate reaction. 'I think we should book the earlier flight to Chicago' is correctly met with 'That works for me — I'll change the reservation now,' not with a fact that happens to share vocabulary with the statement.

Why the correct answer often doesn't sound like an answer

The gap between a mid-range score and a high one on Part 2 is almost entirely about tolerance for indirectness. ETS writes correct responses that deliberately avoid restating the grammar of the question: a yes/no question gets answered with a hedged, real-sounding line like 'Not yet, we're still finalizing the mailing list' instead of a bare 'no'; an embedded question gets answered as if the frame ('Do you know...', 'Could you tell me...') wasn't even there. The classic higher-band pattern is a response that deflects or reframes rather than answering head-on — 'When does the meeting start?' met with 'Didn't you get the email?' — and it is correct precisely because it's the kind of thing a real colleague would say. If you're listening for an answer shaped exactly like the question, you will systematically reject correct options at the harder end of the test.

The trap taxonomy in every wrong option

  • Word-echo traps reuse a word or its root from the question on an unrelated topic. Asked 'Where did you send the client's shipment?', a trap option answers 'We shipped it yesterday afternoon' — reusing 'shipment' as 'shipped' while actually answering 'when', not 'where'.
  • Sound-alike traps swap in a near-homophone: 'back up the server' answered by 'she's not a backup singer anymore', or 'the new assembly line' answered by 'we assembled it ourselves', or 'still valid' answered by 'hasn't validated her parking'. The sound survives the trip from question to option; the meaning doesn't.
  • Wrong-question-type traps answer a real question, just not the one asked — 'Who's interviewing the marketing candidate?' met with 'In the third-floor conference room' (that answers 'where') or 'Sometime around three o'clock' (that answers 'when'). These are the easiest traps to eliminate once you've locked the question word, and the hardest to catch if you tuned in late.
  • Plausible-but-wrong-topic traps are true-sounding statements that simply answer nothing. Asked the asking price for an office space, 'The realtor showed it to three clients' is a believable real-estate fact — it just isn't a price.

The pressure of zero visual support

Compare Part 2 to its neighbors. Part 1 gives you a photograph to anchor meaning before the audio even starts, so a missed word can often be recovered from what you can see. Parts 3 and 4 print the questions in advance and give you a run of several seconds of context — a whole conversation or talk — to triangulate the answer even if one detail slips past. Part 2 gives you nothing to fall back on: no photo, no printed question, no second data point. You get one pass at the question and three passes at responses deliberately built to sound plausible in isolation. Losing focus for the first two words of the question is often unrecoverable, because those two words are what tell you which of the three options is even structurally possible.

Listen for shape, not content — and don't try to take notes

There's no time to write anything down: options arrive only a couple of seconds apart with nothing on screen to reference afterward, so notes are a liability, not a safety net. The workable technique is to hold the question's opening word (or its auxiliary, for yes/no and tag questions) as a single mental filter — a 'shape' the correct answer must fit — and run each of the three options past that filter the instant you hear it, discarding on contact rather than waiting to compare all three at the end. By the time option C finishes, you should already have eliminated two of the three; the pause before the next item is for confirming your instinct, not for reconstructing the question from memory.