- Home
- Methodology: how this test is built and scored
Methodology: how this test is built and scored
Updated October 9, 2026
This page explains exactly how the test works, so you can judge the result for yourself. In short: original puzzles generated by our own code, a fixed blueprint of questions, a simple scoring formula, and an honest estimated range instead of a single number.
The puzzles
Every puzzle is generated by our own software from rules. Nothing is copied from published tests, books or other websites. There are three families:
- Matrices: a 3×3 grid of figures with the last cell missing. Rules such as “one more shape per cell” or “each row uses the same three fills” run along the rows. Six options.
- Number series: five or six numbers following a rule; pick the next one from five options.
- Rotation: a shape made of 5 to 7 squares; pick the option that is the same shape turned, not mirrored or altered. Five options.
Each family has six difficulty levels, defined by the rules involved: for matrices, how many rules apply at once; for series, how complex the rule is; for rotation, the size of the shape and the angles.
Quality checks
A generated puzzle is kept only if it passes automatic checks. Failing puzzles are discarded, never fixed by hand.
- Exactly one answer. For number series, every known rule is fitted to the numbers shown. If two rules fit and predict different next numbers, the puzzle is thrown away. For matrices, the first two rows must determine every rule. For rotation, exactly one option may be a turned copy of the target.
- Wrong answers come from real mistakes. Each wrong option is built from a named error, such as ignoring a rule, applying a step twice, choosing a mirror image or slipping by one.
- No guessing from the options. For matrices, a program that sees only the six options and picks the most “typical” one must not find the answer more often than chance. Our options are balanced so that it never does. For number series, the answer’s position by size among the options is random.
- No colour cues. Figures differ by shape, number, size, fill pattern, turn and marks, never by colour alone.
The puzzles are stored in a versioned item bank with up to 120 distinct puzzles per family and level.
The test
| Quick | Full | |
|---|---|---|
| Scored questions | 16 | 32 (the Quick 16 plus 16 more) |
| Time limit | 12 minutes | 25 minutes |
| Practice | one unscored puzzle per family, with feedback | same |
Questions come in three parts: matrices (7), number series (5) and rotation (4), with difficulty rising inside each part. For each attempt the questions are drawn fresh from the bank, skipping puzzles this browser has already seen. You can go back, skip and change answers. Unanswered questions count as wrong; guessing has no penalty. Speed earns no points. Time is used only to detect attempts too fast to score.
Scoring
- Raw score = the number of correct answers.
- z = (raw score − the average raw score) ÷ the standard deviation of raw scores, both taken from the norms below.
- Estimate = 100 + 15 × z, the standard IQ scale.
- Range = estimate ± the half-width below, kept within the floor and ceiling the test can measure.
The current norms:
Norms version: v0-provisional
| Form | Items | Average raw score | SD of raw scores | Half-width | Measurable range |
|---|---|---|---|---|---|
| Quick | 16 | 9.3 | 3.3 | ±13 | 70–135 |
| Full | 32 | 18.7 | 6.4 | ±10 | 65–140 |
A short test cannot rank the extremes: at the ceiling the result reads “135 or higher”, at the floor “70 or lower”.
If more than a third of the answers took under three seconds each, or the whole test took under two minutes, no range is shown. The attempt was too fast to score reliably.
A second or later attempt in the same browser is labelled practice-affected, because retaking a test raises scores by about five points on average.
Limits
- The norms are provisional. Version v0-provisional is a set of design targets, worked out from the share of people expected to solve each level, not from measurements. It will be replaced once enough anonymous results have been collected.
- Scores are relative to people who took this test, who are not a random sample of the population.
- A short test is imprecise. Sixteen questions place a score within about ±13 points; thirty-two within about ±10.
- It measures a slice of reasoning: visual and numerical pattern-finding, plus spatial rotation. It does not measure knowledge, creativity, language or the many other abilities a full assessment covers.
- It is for curiosity and entertainment. It is not a clinical assessment and must not be used for medical, educational or employment decisions.
Data
No personal data is collected. The site sends a small anonymous record of each finished attempt (which puzzles, right or wrong, how long, form, language, device type) to calibrate difficulty. It never stores names, email addresses, IP addresses or any identifier. See the privacy policy.