Are IQ Tests Biased? Six Questions Hiding Inside One
- Written by
- IQCognify Editorial Team
- Reviewed for accuracy
- IQCognify Research Review Process
- Last updated
- Reading time
- 11 min
Quick answer
Ask whether IQ tests are biased and you will get a confident answer in either direction, usually within a sentence. Both confident answers are premature, and for the same reason: bias is not one question. In psychometrics it names at least six distinguishable things, and a test can be clean on one and fail on another. This guide separates them, reports what the research actually found, and marks the edges of each finding. It does not end with a verdict, because the evidence does not support one.
What does “biased IQ test” mean?
In everyday use, calling a test biased usually means it is unfair to some group, often because its content assumes a particular upbringing.
In psychometrics the word is narrower and technical. A test is biased when people of the same underlying ability get systematically different scores for reasons unrelated to the ability being measured. The emphasis matters: the comparison is between people who are genuinely equal on the thing the test is supposed to measure.
That definition immediately splits into separate empirical questions.
| The question | What it asks | The technical name |
|---|---|---|
| Do scores mean the same thing in both groups? | Whether the measurement model holds across groups | Measurement invariance |
| Does a particular question behave differently? | Whether one item disadvantages equal-ability members of a group | Differential item functioning (DIF) |
| Does the test predict outcomes equally well? | Whether it over- or under-predicts a later outcome for a group | Predictive bias |
| Do the groups average differently? | A descriptive fact about scores | Group mean differences |
| Is using the test here justifiable? | A judgement about values, consequences and context | Fairness |
The first three are statistical, testable, and can come out differently from one another. The fourth is often mistaken for bias and is not. The fifth is not settled by statistics at all.
Measurement invariance
Measurement invariance asks whether a test measures the same construct in the same way across groups. It is tested in tiers, each more demanding than the last: configural invariance means the same general structure holds in both groups; metric invariance means the items relate to the underlying ability with the same strength; scalar invariance means the items also have the same baselines.
Scalar invariance is the level that matters for comparing averages. Without it, a difference in mean scores cannot be cleanly interpreted as a difference in the underlying ability, because part of the gap may come from the measurement rather than from the thing being measured.
Jan Wicherts, writing in The Clinical Neuropsychologist in 2016, argues that this is a core issue in deciding whether population-based norms are valid for subgroups. The paper is a selective rather than systematic review: it discusses why invariance might fail and reviews roughly a dozen studies of invariance in commonly used neurocognitive batteries. In over half of those reviewed studies, the batteries were not found to be measurement invariant across groups based on ethnicity, gender, educational background, cohort or age — and apart from age and cohort, test manuals do not take such lack of invariance into account when computing full-scale IQ scores or normed domain scores.
Read that carefully, because it is easy to over-read
“Over half” is a count of the studies gathered in that review, not an estimate of how often invariance fails across all testing everywhere. The review covers a particular set of batteries and a particular set of groupings, it does not report how large the violations were, and it is not a finding that IQ tests are biased. What it supports is narrower and still substantial: non-invariance is common enough in the reviewed literature to matter, and manuals mostly do not adjust for it.
Differential item functioning
Where invariance is a property of the whole instrument, differential item functioning is about individual questions. An item shows DIF when two people of equal ability but different group membership have systematically different chances of answering it correctly. It is the most concrete form of test bias, because it points at a specific question you can inspect and, if warranted, remove.
Lúcio and colleagues, publishing in Assessment in 2019, examined Raven's Colored Progressive Matrices in 582 Brazilian preschool children — mean age 57 months, 46% female — testing for DIF by sex and age. A few items presented DIF: two for sex and one for age. The authors concluded that the Raven's items were mostly measurement invariant.
Three items out of a reduced form is a small amount of item-level bias, and it is worth stating because it runs against the assumption that non-verbal reasoning tests must be riddled with it.
The scope is equally worth stating
That study looked at Brazilian preschoolers, and at sex and age — not at culture, ethnicity or language background. It is not evidence that Raven's is culture-fair, and we will not describe it that way. Raven's was designed to reduce verbal and cultural loading; designing for something is not the same as demonstrating it, and this study did not test that question. See Raven's Progressive Matrices for what the instrument does and does not do.
Predictive bias
Predictive bias asks a different question: when a test is used to forecast something — job performance, grades — does it forecast equally well for everyone? A test shows predictive bias if it systematically over-predicts or under-predicts the outcome for one group.
Note the shift. This is no longer about whether the score means the same thing. It is about whether the score works the same way as a prediction. A test could be clean here and still fail invariance, or the reverse.
Most of this evidence comes from personnel selection — hiring and job performance. That context should be kept in view, because it is not the same as clinical, educational or self-assessment use, and results from one setting do not automatically transfer to another.
Within that context, Berry and Zhao, writing in the Journal of Applied Psychology in 2015, developed a method intended to avoid a known statistical problem in earlier tests of over- and under-prediction. They reported that African American job performance was typically over-predicted by cognitive ability tests across levels of job complexity. Over-prediction means the test forecast higher performance than was subsequently observed — the opposite direction from what a bias-against hypothesis would predict.
Group differences are not automatically measurement bias
This is the most common error in public discussion, so it gets its own section.
If two groups have different average scores, that is a descriptive fact about the scores. It is not, by itself, evidence that the test is biased. The difference might reflect a real difference in the measured ability, or unequal access to whatever the test draws on, or measurement problems — and telling those apart is exactly what invariance and DIF analyses are for.
The distinction is standard in the field rather than a rhetorical move. The 2023 analysis discussed in the next section notes explicitly that smaller mean differences between groups do not establish fairness: the two questions are separate, and a test with small gaps can still predict unequally while a test with large gaps might not.
Running the inference in either direction — the groups differ, so the test is biased; or the test is unbiased, so the gap is real — skips the step that would settle it.
What research actually finds
Pulling the threads together, scoped as each source scopes itself.
- On invariance: in a selective review of roughly a dozen studies of common neurocognitive batteries, over half failed invariance on at least one grouping, and manuals largely do not adjust for this except by age and cohort (Wicherts, 2016).
- On DIF: on Raven's Colored Progressive Matrices, in 582 Brazilian preschoolers, tested by sex and age, item-level bias was minimal — three items (Lúcio and colleagues, 2019).
- On predictive bias in hiring: cognitive ability tests were found to over-predict African American job performance across job complexity levels (Berry and Zhao, 2015).
- On predictive bias, contested: for Hispanic test takers in personnel selection, one analysis reported under-prediction and a later re-analysis reported over-prediction.
- On fairness: a distinct question that statistical results inform but do not resolve.
No single verdict follows from that list. What follows is that the answer depends on which question you asked, about which test, for which groups, and for which use.
Why conclusions can differ between studies
The clearest illustration comes from two analyses of the same question about the same group, both in personnel selection.
Berry and colleagues reported in 2020 that cognitive tests under-predict Hispanic job performance by around 0.21 standard deviations — a result that would count against the fairness of using them in selection.
Sackett, Zhang and Berry revisited this in the Journal of Applied Psychology in 2023. They located 119 studies in which all three relevant quantities could be drawn from the same sample and setting, rather than combined across different sources, and reached the opposite conclusion: that tests over-predict Hispanic performance by between 0.04 and 0.20 standard deviations, depending on assumptions made about artifact corrections. They attribute the difference to specific method choices — how range restriction was corrected, how comparable the samples were, and which subgroup the validity estimate came from.
Two things are worth taking from this, and they pull in different directions. First, predictive-bias conclusions are genuinely sensitive to methodology: a headline result can reverse under different but defensible analytic choices, so anyone citing a single study as settling the matter is overstating what the literature supports. Second, this is not a reason to dismiss the research. It is what a field looks like when it is working — a published claim, a documented re-analysis, and an explicit account of which choices mattered. Berry appears as an author on both sides of the exchange, which is not the signature of a field defending a position.
What this means for interpreting IQ scores
A score is tied to the group it was normed on. Norms describe a standardisation sample, and when someone differs substantially from that sample the comparison rests on weaker ground. Invariance research is how that concern gets tested rather than merely asserted. The mechanics of norming are covered in how IQ tests work.
“Is this test biased?” is therefore incomplete as a question. Biased in which sense, for which groups, for which use? A test can be adequate for one purpose and poorly suited to another.
Test-level claims rarely transfer, either. Findings about one battery, in one population, on one grouping do not automatically hold elsewhere — and that cuts both ways. Reassuring findings do not generalise any more readily than alarming ones.
Finally, the stakes set the standard. For a free online test that tells you something about yourself, the bar is honesty about limits. For hiring, clinical or educational decisions, the bar is a properly normed instrument with published evidence about the groups it will be used on.
What IQCognify has and has not tested
Given everything above, it would be inconsistent to leave our own position vague. IQCognify's test has not been examined for measurement invariance across any grouping, for differential item functioning on any item, for predictive bias against any outcome, or for subgroup fairness of any kind.
No analysis, no claim
No such analysis has been conducted or published by us, and nothing on this site should be read as implying otherwise.
Two consequences follow honestly. Our items were written in English and draw on conventions — vocabulary, formats, the assumption that abstract puzzles are a normal thing to be shown — that are not evenly distributed, and we have not measured how much that matters. And a short online test is not a sound basis for any decision about another person, for these reasons alongside the ones set out in are online IQ tests accurate.
We would rather state the gap than imply a clean bill of health we have not earned.
Sources
Each source below is cited only for what it reports, in the population and setting it studied. Except where noted, these were reviewed at abstract level rather than in full text.
- Wicherts, J. M. (2016). The importance of measurement invariance in neurocognitive ability testing. The Clinical Neuropsychologist, 30(7), 1006–1016. doi:10.1080/13854046.2016.1205136. A selective review of roughly a dozen invariance studies.
- Lúcio, P. S., Cogo-Moreira, H., Puglisi, M., Polanczyk, G. V., & Little, T. D. (2019). Psychometric investigation of the Raven's Colored Progressive Matrices test in a sample of preschool children. Assessment, 26(7), 1399–1408. doi:10.1177/1073191117740205. 582 Brazilian preschoolers; DIF tested by sex and age only.
- Berry, C. M., & Zhao, P. (2015). Addressing criticisms of existing predictive bias research: cognitive ability test scores still overpredict African Americans' job performance. Journal of Applied Psychology, 100(1), 162–179. doi:10.1037/a0037615. Personnel-selection context.
- Sackett, P. R., Zhang, C., & Berry, C. M. (2023). Challenging conclusions about predictive bias against Hispanic test takers in personnel selection. Journal of Applied Psychology, 108(2), 341–349. doi:10.1037/apl0000978. Re-analysis of 119 studies; personnel-selection context.
- Holden, L. R., & Tanenbaum, G. J. (2023). Modern assessments of intelligence must be fair and equitable. Journal of Intelligence, 11(6), 126. doi:10.3390/jintelligence11060126. Cited for fairness as a distinct concern; the authors argue for a particular theoretical position.
Berry and colleagues (2020) is described only as characterised by Sackett, Zhang and Berry (2023), which we read at abstract level; we did not retrieve the 2020 paper directly.
Frequently asked questions
Are IQ tests biased?+
It depends which sense you mean. Bias has several distinct technical meanings and the research answers them differently: measurement invariance frequently failed in the studies gathered by one selective review, item-level bias on the one test examined here was minimal, and predictive-bias findings in hiring are contested. There is no single yes or no that the evidence supports.
What is measurement invariance?+
Whether a test measures the same thing, in the same way, across groups. Scalar invariance — the strictest commonly tested level — is what is needed before average scores can be compared cleanly.
What is differential item functioning?+
When a specific question behaves differently for equal-ability members of two groups. It is the most concrete form of test bias, because it identifies a particular item that can be inspected and, if warranted, removed.
Does a difference in average scores prove bias?+
No. It is a descriptive fact about scores. Whether it reflects the measured ability, unequal circumstances, or a measurement problem is a separate question that invariance and DIF analyses are designed to address.
Is Raven's Progressive Matrices culture-fair?+
We would not put it that way. Raven's was designed to reduce verbal and cultural loading, and the study reviewed here found little item-level bias by sex and age among Brazilian preschoolers — but that study did not test cultural groups. Designing for a property is not the same as demonstrating it.
Has IQCognify's test been checked for bias?+
No. It has not been examined for measurement invariance, differential item functioning, predictive bias or subgroup fairness, and we do not claim otherwise.
Find out your IQ
Take the free IQ test and get your score, percentile, and a full cognitive breakdown in about 12 minutes.
Start Free Test