How IQ Tests Work: The Ultimate Plain-English Guide (2026)

How IQ tests work, from Alfred Binet's 1905 original to modern item banks: the four cognitive domains, how raw answers become a score, and what the g-factor actually means.
Picture of Mindaura

Mindaura

Leading platform for validated cognitive and personality assessments.

how IQ tests work

Understanding how IQ tests work starts with one idea: a good test never asks what you already know, it asks how efficiently you reason with something new. That single design principle, more than any specific question format, is what separates a real cognitive assessment from a trivia quiz wearing an IQ test’s name.

This guide walks through how IQ tests work in plain English: where the format came from, the four cognitive domains almost every test samples, exactly how raw answers become a comparable score, what the g-factor is and why it matters, and the specific misconceptions that trip people up when they think about testing methodology.

🧠 IQ & Aptitude Test

How sharp is your mind?

Take the Quiz Test and get your personalised report instantly.

Start the Test →

⏱ 5 min  ·  📊 Instant score

A Brief History of How IQ Tests Work

The story of how IQ tests work today starts in 1905, when French psychologist Alfred Binet built the first practical intelligence test to identify schoolchildren who needed extra academic support, not to rank the gifted but to catch struggling students early. Binet’s test used age-graded tasks: a child’s performance was compared to what was typical for their age.

Lewis Terman adapted Binet’s scale for American use in 1916 as the Stanford-Binet, introducing the literal intelligence quotient formula: mental age divided by chronological age, multiplied by 100. That’s genuinely where the word “quotient” comes from. A child with a mental age of 12 and a chronological age of 10 scored 120 under this original system.

In 1939, David Wechsler changed how IQ tests work at a structural level by replacing that ratio with a deviation-based score built on a normal distribution, with a mean of 100 and a standard deviation of 15. That is the method still used by the WAIS-IV and virtually every modern scale today. The ratio method broke down badly for adults, since mental age plateaus while chronological age keeps climbing; Wechsler’s deviation method fixed that by comparing each person only to same-age peers.

Modern test construction now draws on Cattell-Horn-Carroll (CHC) theory, which organizes cognitive abilities into a hierarchy rather than treating intelligence as one flat skill. The field has also had a troubling history of misuse, including its role in early-20th-century eugenics policy; modern test developers and psychologists are explicit that a score describes a specific set of cognitive skills, not a person’s worth.

The Four Cognitive Domains Every Test Measures

However a specific test is branded, how IQ tests work under the hood almost always comes down to sampling performance across a handful of cognitive domains:

Abstract / Logical Reasoning

Pattern recognition and visual matrices, the strongest marker of fluid intelligence and the domain most resistant to coaching.

Verbal Reasoning

Vocabulary and word relationships, reflecting crystallized knowledge built up through education and experience.

Numerical Reasoning

Number sequences and mathematical logic, testing pattern detection more than arithmetic speed.

Spatial Reasoning

Mental rotation and 3D visualization, measuring how accurately you can manipulate shapes that exist only in your head.

Clinical batteries like the WAIS-IV group these into four index scores: Verbal Comprehension, Perceptual Reasoning, Working Memory, and Processing Speed. These combine into the Full Scale IQ. Free online tests typically compress this into fewer, shorter sections, which is a reasonable tradeoff for accessibility but means less precision per domain than a full clinical battery.

How Scoring Actually Works

Once the questions are answered, how IQ tests work moves from item-level responses to a single comparable number in a defined sequence of steps, not a single calculation:

  1. Raw score: the simple count of correct answers, sometimes weighted by item difficulty.
  2. Normative comparison: that raw score is measured against a large validation sample of people who already took the same test, usually stratified by age and sometimes by education level.
  3. IQ-scale conversion: the comparison is converted onto the standard scale, where 100 marks the average and each 15-point step marks one standard deviation.
  4. Domain subscores: a full profile is generated to show relative strengths across domains, not just one headline number.

A score of 115, for instance, means performing better than roughly 84% of the norm group. That is a concrete, population-anchored statement, not a vague label. Building that norm group is itself a large undertaking: publishers recruit thousands of people across age brackets, regions, and demographics specifically so the resulting comparison is representative rather than skewed toward whoever happened to be easiest to test.

The g-Factor: What These Tests Are Really Trying to Capture

In 1904, psychologist Charles Spearman noticed something that still shapes how IQ tests work over a century later: people who score well on one cognitive task also tend to score well on unrelated ones. He called the common thread running through all of them the g-factor, or general intelligence.

Under CHC theory, g sits at the top of a three-level hierarchy. Below it sit broad abilities: fluid intelligence (Gf), crystallized intelligence (Gc), visual processing (Gv), working memory (Gwm), and processing speed (Gs). Below those sit dozens of narrower, specific skills. Fluid intelligence, solving novel problems with no reliance on prior learning, correlates most closely with g, which is exactly why abstract matrix problems carry so much weight in most modern batteries. For the research literature behind this idea, see the Stanford Encyclopedia of Philosophy’s overview of psychometric intelligence.

Types of IQ Tests

Not every instrument that measures intelligence looks the same, and how IQ tests work in practice depends partly on which format is being used. A few structural distinctions matter most:

  • Individual vs group administration: a one-on-one clinical test allows a trained examiner to observe effort and catch irregularities in real time; group tests trade that observational detail for scale and speed, and are common in educational or occupational screening.
  • Verbal vs nonverbal / culture-reduced: nonverbal tests like Raven’s Progressive Matrices minimize language and cultural knowledge, which matters for testing across different linguistic or educational backgrounds, though no test is entirely free of cultural influence.
  • Power vs speeded tests: a power test allows generous time and measures how far reasoning holds up on hard items; a speeded test measures how much you can complete quickly on easier items. Most modern batteries blend both formats across different subtests to capture both dimensions.
  • Clinical vs free online: both can use validated item content, but only a clinical administration adds a trained examiner, controlled conditions, and formal interpretation of results against a documented normative sample.

Reliability and Validity: How Test Quality Is Actually Measured

Two separate properties determine whether a test is well-built, and conflating them is a common source of confusion about how IQ tests work. Reliability is consistency: does the same person get a similar score if retested under similar conditions? Validity is relevance: does the score actually support the conclusion being drawn from it, whether that’s general cognitive ability, academic readiness, or something else?

A test can be highly reliable but poorly valid for a specific purpose, consistently measuring something real, just not the thing you think it’s measuring. Clinical scales like the WAIS-IV report test-retest reliability figures typically above 0.90, alongside decades of validity research tying scores to academic and occupational outcomes. Free tests vary much more in how rigorously they report either figure, which is part of why methodology transparency matters when evaluating one; see are online IQ tests accurate? for the deeper dive on this distinction.

Who Administers and Interprets These Tests

A clinical IQ assessment is administered and interpreted by a licensed psychologist trained specifically in psychometric assessment, not a general practitioner or an untrained proctor. That training matters for two reasons: standardized administration, where every test-taker gets the same instructions and timing so scores stay comparable, and clinical judgment, where an examiner can flag when a low score reflects anxiety, fatigue, or a language barrier rather than genuine ability. Free online tests, by construction, can standardize the questions but not the examiner’s judgment. That is one more reason a self-administered score is a useful estimate rather than a diagnostic instrument.

Item Banks and Validation

Not every test that claims to measure IQ has been through the same rigor. Reputable instruments, whether clinical or free, draw from validated item banks, question sets that have themselves been checked against known cognitive measures. The Young & Keith (2020) validation study is a good example: it checked an ICAR-based public item bank against the WAIS-IV and found correlations of r = 0.70 to 0.85, evidence that a well-constructed free test can track the same underlying construct as a full clinical battery, even without the same precision. For more on what that accuracy gap looks like in practice, see are online IQ tests accurate?

A Worked Example: From Raw Score to IQ Score

Abstract steps are easier to follow with numbers attached. Say a test-taker answers 27 out of 33 items correctly on an abstract reasoning section. That raw score of 27 means nothing on its own. It only becomes meaningful once it’s placed against the norm sample’s distribution of raw scores for that exact item set.

If the norm sample shows a mean raw score of 22 with a standard deviation of 5 raw points, a raw score of 27 sits exactly one standard deviation above that mean. Converted onto the IQ scale, one standard deviation above the mean is 115. That 115 is now comparable to anyone else’s 115 on this test, regardless of how many raw items existed or how hard they were. The conversion step is what makes cross-test and cross-person comparison possible at all.

How Test Publishers Build and Validate New Items

New items don’t appear on a live test until they’ve been through a pilot phase: administered to a trial sample, checked for how well they discriminate between higher and lower scorers, and screened for unintended bias across demographic groups. Items that are too easy, too hard, ambiguous, or that show unusual score gaps between groups matched on ability get revised or dropped before the test ever reaches a real test-taker. This is also how test publishers periodically refresh norms to account for the Flynn Effect, the gradual population-level drift in average raw scores over time.

Common Misconceptions About How IQ Tests Work

A few myths persist even though they run against how these instruments are actually built:

  • “It’s full of trick questions.” Validated items are piloted and revised specifically to remove ambiguity; a well-built test rewards clear reasoning, not lateral thinking about wording.
  • “A truly fair test would be culture-free.” No test is entirely free of cultural influence, but nonverbal and abstract-pattern formats meaningfully reduce that dependence compared to vocabulary-heavy formats.
  • “Faster answers mean higher intelligence.” Only on speeded subtests is raw speed the point; on power-format sections, working through a hard item correctly matters more than finishing quickly.
  • “One test measures your fixed, permanent intelligence.” Scores shift with practice effects, test conditions, education, and age. See average IQ by age for how much this varies across the lifespan.

What This Process Does Not Measure

Knowing how IQ tests work also means being honest about their limits. By design, they exclude creativity and divergent thinking, emotional intelligence, practical wisdom and street smarts, persistence and self-discipline, and moral character and social skill.

A test built this way is deliberately narrow. That narrowness is a feature, not a flaw. It is what makes the resulting number comparable across people at all, but it also means a single score should never stand in for a full picture of someone’s abilities. For how to weigh a result against those limits, see how to interpret your IQ score.

How This Compares to a Free Online Test

Everything above describes methodology, not delivery format. A well-built free online test can use the exact same domains, item-bank logic, and deviation scoring as a clinical instrument. The difference lies in supervision, standardized conditions, and a trained examiner’s ability to catch irregularities a self-administered format can’t. See our free IQ test with instant results for how Mindaura applies this same methodology.

Common Questions

Who invented the first IQ test?

Alfred Binet, in 1905, to identify French schoolchildren who needed extra academic support.

What is the g-factor?

General intelligence, the statistical tendency for performance across different cognitive tasks to correlate, first identified by Charles Spearman in 1904.

Do all IQ tests measure the same thing?

Not exactly. They vary in item format and scale, but validated tests are built to converge on the same underlying construct, general cognitive ability.

Why do abstract reasoning questions matter so much?

Because they measure fluid intelligence with the least reliance on prior learning, making them the strongest single marker of general cognitive ability.

How long does a clinical IQ test take?

A full WAIS-IV administration typically runs 60 to 90 minutes, one-on-one with a trained examiner.

What’s the difference between a power test and a speeded test?

A power test gives generous time and measures how far reasoning holds up on hard items; a speeded test measures how much you can complete quickly on easier ones. Most modern batteries use both formats across different subtests.

Can practice change how I score on a test like this?

Repeated exposure to similar items produces a practice effect, a score bump from familiarity rather than a genuine change in underlying ability. That’s why a first serious attempt is usually the most meaningful one.

How are new test items checked for bias?

Publishers pilot new items on trial samples and screen for unusual score gaps between demographic groups matched on overall ability, dropping or revising items that show unintended bias before they appear on a live test.

Does how IQ tests work differ for children versus adults?

The core methodology is the same, but children’s batteries like the WISC-V use age-graded item sets and norm groups, since cognitive development changes rapidly across childhood in a way adult norms don’t need to account for.

Is a higher raw score always a higher IQ score?

Not directly. A raw score only becomes an IQ score once it’s converted against the norm sample’s distribution, so the same raw score can convert to different IQ scores on different tests depending on how difficult that test’s item set is.

Final Thoughts

How IQ tests work comes down to a simple chain: sample reasoning across a few cognitive domains, compare the result to a large, carefully validated normative population, and convert that comparison into a score built on a normal distribution. Every step in that chain, from item piloting to norm-group construction to the final deviation score, exists to make one person’s number meaningfully comparable to everyone else’s.

Curious how you’d score using this exact methodology? Take the Mindaura free IQ test and see your results across all four domains instantly.

Related Reading

🧠 IQ & Aptitude Test

How sharp is your mind?

Take the Quiz Test and get your personalised report instantly.

Start the Test →

⏱ 5 min  ·  📊 Instant score

Picture of Alex

Alex

Alex writes about cognitive assessment, IQ testing methodology, and how intelligence is measured and interpreted. His work focuses on making psychometric concepts clear and accurate, covering topics like fluid vs. crystallized intelligence, how IQ scores are calculated, and what test results do and don't tell you about a person's abilities.

In this article

🧠 IQ & Aptitude Test

How sharp is your mind?

Take the Quiz Test and get your personalised report instantly.

Start the Test →

⏱ 5 min  ·  📊 Instant score

Keep reading

The Mensa IQ test guide: the 98th-percentile qualifying threshold, matching...
15 free practice IQ questions across abstract, verbal, numerical, and...
Can you improve your IQ? What the evidence shows about...