The whole exam on Duolingo's own timings, with every item chosen by how you are doing rather than dealt from a fixed set. Reading and listening come back marked item by item; writing and speaking come back rated, and the report says which is which.
13 task types · scored 10–160 · ~60 minutes
Six invented responses put through the real estimator, so these numbers cannot disagree with the product. The estimate follows the difficulty you sustain rather than the count you get right, and the range narrows as the evidence arrives.
A one-parameter logistic ability model reads how you are doing and picks the next item at the difficulty that measures you best — which is how the real test works and where its score comes from. Practice material is almost always a fixed set of questions.
Duolingo changed the format on 1 July 2025: Read Aloud and Listen, Then Speak were removed and Interactive Speaking was added. All thirteen current task types are here, and the deleted two are not — a lot of material still ships them.
Reading and listening are marked against answer keys, so you get the item-by-item record. Writing and speaking are rated from your own text and audio. The report tells you which number is which instead of printing one confident total.
Timings are Duolingo's, per task, and they are hard — the real test cuts you off and so does this. Counts per sitting sit inside the published range for every task, and they are fixed rather than randomised so two of your own attempts are comparable.
| Task | On the clock | Per sitting | Feeds | Item choice |
|---|---|---|---|---|
Read and Select Decide whether each word is real English. Five seconds each. | 5s | 18 | Reading | Adaptive |
Fill in the Blanks Complete the missing letters of one word in a sentence. | 20s | 6 | Reading | Adaptive |
Read and Complete Restore the missing letters across a damaged passage. | 3:00 | 4 | Reading | Adaptive |
Interactive Reading One passage, six questions, one shared clock. | 7:30 | 2 | Reading | Adaptive |
Listen and Type Type the sentence you hear. Up to three plays. | 1:00 | 8 | Listening | Adaptive |
Interactive Listening Follow a conversation, choose replies, then summarise what happened. | 6:30 then 1:15 | 2 | Listening + Writing | Adaptive |
Write About the Photo Describe the picture in one minute. | 1:00 | 3 | Writing | Fixed |
Interactive Writing Write for five minutes, then extend your answer for three more. | 30s prep + 5:00 then 3:00 | 1 | Writing | Fixed |
Speak About the Photo Describe the picture aloud for up to ninety seconds. | 20s prep + 1:30 | 1 | Speaking | Fixed |
Read, Then Speak Read a prompt, then speak to it for up to ninety seconds. | 20s prep + 1:30 | 1 | Speaking | Fixed |
Interactive Speaking A conversation. Thirty-five seconds per answer, no second takes. | 35s | 1 set · 7 questions | Speaking | Adaptive |
Writing Sample Five minutes. Scored, and sent to the institutions you apply to. | 30s prep + 5:00 | 1 | Writing | Fixed |
Speaking Sample Up to three minutes. Shared as video with your institutions. | 30s prep + 3:00 | 1 | Speaking | Fixed |
The extended responses are marked “Fixed” because their difficulty is not what the test manipulates: there is one prompt and one clock, and what varies is your answer.
The overall is the mean of the four individual subscores, rounded to the nearest five, which is Duolingo's own arithmetic. It has a consequence candidates consistently get wrong: one weak skill drags the overall down by a quarter of its shortfall, and strength elsewhere does not compensate. Every subscore is reported with the range it plausibly sits in and the number of responses behind it.
A subscore with nothing behind it is left empty rather than filled in, and the overall is withheld until all four exist.
Most of this test has a right answer, and a smaller part of it needs somebody to form an opinion. Keeping the two apart is the whole reason the report can be trusted: a tool that sends the sitting to a model and prints whatever comes back cannot tell you which of its numbers is a fact.
Marked against answer keys that travel with the material. The same responses score identically every time, and the report shows the record: which word you rejected, which gap you missed, which of the dictation's words arrived and which did not.
Your own text and your own audio, judged against CEFR criteria by a model — sampled several times and reduced to a median rather than reported from one pass. Where the passes disagree by more than half a level, the report says so.
It is not a Duolingo score and it does not predict one. What the engine measures is which difficulty of item you could sustain, which is honest arithmetic over your actual responses. Converting that to a number on the 10–160scale is our own calibration against Duolingo's published CEFR bands, because Duolingo's conversion is not public and its item difficulties come from millions of sittings while ours come from the judgement of the people who wrote the items. Treat the number as a well-founded estimate of where you stand.
A 105 means nothing to somebody who has only ever thought in IELTS bands, so the report maps your score to its CEFR level and to the rough equivalents on the tests you are choosing between.
| CEFR | Duolingo scale | IELTS, approximately | TOEFL iBT, approximately |
|---|---|---|---|
| A1–A2 | 10–55 | 0–4.0 | 0–17 |
| B1 | 60–95 | 4.0–5.5 | 18–64 |
| B2 | 100–125 | 5.5–6.5 | 65–97 |
| C1 | 130–150 | 7.0–8.0 | 98–118 |
| C2 | 155–160 | 8.0–9.0 | 119–120 |
Concordances are estimates for orientation, not admission equivalences. Every institution sets its own requirement, and the CEFR bands are the boundaries Duolingo publishes rather than something interpolated here.
Standing is judged against your own mean rather than an absolute bar, because somebody scoring well across the board does not need to be told their weakest task is bad — they need to know which one is holding the average down, which on this test is exactly what decides the overall. The advice is fixed and specific to how the task is marked, so it is the same for everyone with the same profile and you can see why it was given.
An omitted word is penalised more heavily than a misspelling, so type a guess for anything you did not catch — a wrong guess costs no more than the blank, and a near miss earns partial credit.
The two Highlight questions carry as much as the four gaps, and the whole set shares one clock. Over-highlighting costs marks, so select only the words that answer the question.
Thirty-five seconds leaves no room for a warm-up sentence. Answer in your first clause, add one reason, add one example, stop.
The invented words are built from real stems with endings that do not exist. Read the ending rather than the shape of the word, and decide inside five seconds.
Counted from the banks rather than claimed: 1300 items across the 13 tasks, and the number above is how many complete sittings the first bank to run dry can serve. A DET item is a fixed text with a fixed answer, so meeting one twice tests whether you remember our material rather than whether you know English.
| Free practice sets | ChatGPT | MockLive | |
|---|---|---|---|
| All thirteen current task types | — | — | ✓ |
| Each item chosen by your running ability | — | — | ✓ |
| Official per-task timings, nothing pausable | — | — | ✓ |
| Speaking rated from your own audio | — | — | ✓ |
| All nine scores, individual and integrated | — | — | ✓ |
| New material on every sitting | — | ✓ | ✓ |
One price worldwide — $39.90 for 20 full adaptive mocks. Enough to sit one every few days in the two months before your test and watch the weak subscore come up. Your first session is free — no card, nothing to cancel.
Around 60 minutes for most people, and it can run to 85 if you use every clock to the last second. Both numbers are added up from the tasks themselves rather than rounded off: each of the 13 task types keeps the real test's own timing, and the difference between the two figures is simply that nobody spends the full three minutes on a passage they finished in two.
Really adaptive. Every response is converted to a difficulty and an outcome on the same logit scale, your ability is re-estimated with a Rasch model, and the next item is drawn from the ones closest to that estimate. Each of the four skills is tracked separately, because someone who reads well and listens badly has two different abilities and serving them one set of items would measure neither.
It is an estimate, and it is not a prediction. What we can measure is which difficulty of item you sustained; turning that into a number on the 10–160 scale is our own calibration against Duolingo's published CEFR bands, because Duolingo's conversion is not public. The report also shows how wide the plausible range is and how many responses each subscore rests on, which is the part a single number hides.
All 13 of the current ones, including Interactive Reading, Interactive Listening, Interactive Writing and Interactive Speaking. Read Aloud and Listen, Then Speak were removed from the test on 1 July 2025 and are not here, which is worth checking before you practise from anything else.
The individual subscores are Reading, Writing, Listening and Speaking. The integrated ones — Literacy, Comprehension, Conversation, Production — are averages of two of those each, so they are a reshuffle of the same four numbers rather than a second opinion. They matter because institutions set cut-offs on them.
20. The banks hold 1300 items across the 13 tasks, every sitting draws material you have not seen, and the count is taken from the first bank that would run dry rather than an average. The lobby tells you how many fresh sittings you have left.
By a model reading your text and listening to your audio, against CEFR criteria, sampled several times and reduced to a median rather than reported from a single pass. Where a rating does not come back, that subscore is left empty and the overall is withheld, because an average of three skills and a guess at the fourth is worse than no number.
No. Your first session is free, and passes are one-time payments with no subscription and no card kept on file.
One sitting, about 60 minutes, and a report that names the task holding your average down. The first one is free.
Start the free test