How the test works
Mensura follows the same theoretical foundations as modern intelligence tests: the CHC model of abilities, Item Response Theory and adaptive testing. Here is what happens behind each question, and also what the test is not.
What is measured
The Cattell-Horn-Carroll (CHC) model organizes intelligence into a general ability (g) and broad abilities. The test has seven question types, each tied to one of them:
| Question type | Ability (CHC) | Original paradigm |
|---|---|---|
| Matrices | Gf, inductive reasoning | Progressive matrices (Raven) |
| Number series | Gf/Gq, quantitative induction | Number series |
| Logical deduction | Gf, sequential reasoning | Linear syllogisms |
| Balances | Gq, quantitative reasoning | Balance problems |
| Mental rotation | Gv, visual processing | Shepard and Metzler |
| Working memory | Gwm | Visuospatial span and digits (forward and backward) |
| Verbal analogies | Gc, verbal knowledge and reasoning | Analogies A : B :: C : ? |
Every test spreads its questions evenly across the seven types, so the final score reflects a broad set of abilities rather than a single kind of task.
Why no two tests are the same
Questions do not come from a fixed bank. Each one is built on the spot by a generator (automatic item generation) from rules whose difficulty is known. In a matrix, for example, difficulty grows with the number of rules acting at once and with the kind of rule: progressions are easier than distributions of three values. Every question stores a random seed and a signature, so the system never shows the same question twice to the same person.
The answer options follow a method too. In matrices, the eight options are combinations of three attributes, each with either the right value or a plausible wrong one. Every feature appears in half of the options, so the trick of "picking the figure most similar to the others" does not work.
How difficulty adapts to you
The test is adaptive. After each answer, the system re-estimates your ability and picks the next question at the difficulty that tells the most about you at that point. People who answer correctly get harder questions; people who miss get easier ones. The first three are a little easier, as a warm-up. That is why almost everyone gets between half and two thirds right, and that is expected.
How the score is calculated
Each question type and level has parameters from the three-parameter logistic model (3PL) of Item Response Theory:
P(correct | θ) = c + (1 − c) / (1 + e−1.7·a·(θ − b))
- θ is your ability, on the population scale (mean 0, standard deviation 1);
- b is the question's difficulty; a, how well it separates people who know from people who don't;
- c is the chance of guessing right (1 divided by the number of options).
Ability is estimated with the EAP method (expected a posteriori), which combines all answers with a normal prior distribution. The score shown uses the traditional IQ scale:
score = 100 + 15 · θ
The system also computes the standard error of the estimate and shows the 95% confidence interval. It is narrower in longer tests.
Precision
We simulated thousands of virtual participants with known ability answering according to the model. The typical difference between estimated and true score was:
| Questions | Typical error (points) | 95% interval |
|---|---|---|
| 20 | 5.2 | ±9 |
| 40 | 3.7 | ±7 |
| 60 | 3.2 | ±5 |
| 100 | 2.3 | ±4 |
At the extremes (below 70 or above 130) the estimate leans slightly toward the center, a known effect of the EAP method.
Calibration with real data
At first, each question's parameters come from the theoretical difficulty model of its generator. As people answer, a routine re-estimates those parameters from real data and the engine starts using them. There is not enough data for any recalibration yet.
What this test is not
- It is not a clinically validated psychological test. Tests used for diagnosis, reports, hiring or admissions must be approved by the relevant professional bodies (in Brazil, the SATEPSI of the Federal Council of Psychology) and administered by psychologists.
- The reference population is theoretical. Clinical tests are normed on large, representative samples of a population. Here, the scale starts from the model and is refined with the participants themselves, who are not a representative sample.
- Conditions are not controlled. Fatigue, distractions, hurry, a small screen and practice with this kind of question all change the result.
Use the result as a fun, informative snapshot of your reasoning skills. If you need a real assessment, look for a licensed psychologist.
References
- Carpenter, P. A., Just, M. A. & Shell, P. (1990). What one intelligence test measures: a theoretical account of the processing in the Raven Progressive Matrices Test. Psychological Review, 97(3).
- Embretson, S. E. (1998). A cognitive design system approach to generating valid tests. Psychological Methods, 3(3).
- McGrew, K. S. (2009). CHC theory and the human cognitive abilities project. Intelligence, 37(1).
- Shepard, R. N. & Metzler, J. (1971). Mental rotation of three-dimensional objects. Science, 171.
- Wainer, H. (ed.) (2000). Computerized Adaptive Testing: A Primer. Lawrence Erlbaum.