Study on the flyHow we write questions

Study on the flyHow we write questions

How our questions are written and checked

Every practice question and explanation on Study on the fly is written by AI models and checked by other AI models and by automated tests; the short topic introductions are AI-written too. None of it is official exam material. This page explains how it is made, what is checked, what the checks found, and how to tell us about a mistake.

Written to each exam’s own format

Before any question is written, we set out the exam’s public blueprint: its sections, timing, question types and answer-choice conventions. Questions are then written in small batches to that specification. Each must have exactly one defensible answer, and each wrong choice should reflect a mistake a real student makes. Passages and questions are original; nothing is copied from an official test.

Modeled on real papers, which we do not republish

To match each exam’s topics and difficulty, we studied its past official papers: which topics come up, how often, and how hard the questions are. Practice questions and past-paper questions are rated on one difficulty scale by independent AI raters, and our practice tests are assembled to follow the real papers’ topic and difficulty spread. These ratings are estimates.

The past papers themselves are not published here or in the apps. Where a subject page shows how past questions were spread across its topics, it gives counts only.

How each question is checked

  • Automated tests. Every question must have a valid format, exactly one keyed answer and distinct choices. Math must typeset, and a figure must be drawn to the measurements it states unless it is marked as not drawn to scale.
  • A blind solve. A different AI model answers the question without seeing our key. If it disagrees, or finds two defensible answers, the question is fixed or removed.
  • Review rounds over the whole bank. One model flags anything that looks wrong, a second rules on each flag and writes the fix, and the fixes are then reviewed in turn. Any changed answer key is solved again from scratch.

What the checks found

On October 4, 2026, we drew 3,000 questions at random from the question banks of all 10 exams our apps cover, in English and Turkish. Two AI solvers from different model families answered each one without the key, and a third re-checked every question they flagged. Of the 2,999 questions both solvers answered, problems were confirmed in 7: about 0.2% (95% range 0.1% to 0.5%).

This measures only what an AI solver can catch. A subtle error that fools every checker would not show up, so read the figure as a floor, not a guarantee.

What appears on this site

Topic pages show up to three sample questions each, with answers and explanations, chosen by a fixed rule: text-only questions of middle difficulty that carry a written explanation. Each exam also has one complete practice test here (the same paper the app serves as Practice Test 1) and, where one exists, a harder Challenge paper.

Score calculators, and the scores in the apps, are our own estimates. Only the testing organization can give you an official score.

The vocabulary list has its own page: how the word list was built.

Found a mistake?

Errors can remain. If a question, answer or explanation looks wrong, write to [email protected] with the page address and the question number. A confirmed mistake is fixed in the question bank itself, so the fix reaches this site and the apps alike.

Free and ad-free, with nothing to sign up for to read it.

Browse the exams