Last updated, September 09, 2026
AI IQ Test Leaderboard: Mensa vs Offline
| Model | Mensa Norway Score | Offline Test Score |
|---|---|---|
| GPT 5.6 TERRA Ultra (Vision) | 141 | 136 |
| GPT 5.6 TERRA Ultra | 142 | 133 |
| GPT 5.6 SOL Ultra (Vision) | 124 | 132 |
| Claude-5.1 Fable | 151 | 130 |
| Claude-5.1 Fable (Vision) | 145 | 130 |
| GPT 5.6 SOL Ultra | 145 | 130 |
| GPT 5.6 LUNA Max (Vision) | 145 | 130 |
| Claude-5 Opus | 141 | 129 |
| Grok 4.6 High | 144 | 128 |
| Claude-5 Opus MAX | 141 | 128 |
| Gemini 3.7 Flash | 143 | 128 |
| GPT 5.6 LUNA Max | 145 | 127 |
| Kimi-K3-Vision | 139 | 123 |
| Gemini 3.1 Pro | 144 | 122 |
| Qwen 3.8 Max | 142 | 122 |
| Gemini 3.1 Pro (Vision) | 128 | 119 |
| Kimi K2.6 | 136 | 117 |
| Claude-5 Opus (Vision) | 139 | 116 |
| MiMo V2.5 Pro | 127 | 116 |
| Claude-5 Sonnet | 137 | 113 |
| Grok 4.6 High (Vision) | 131 | 112 |
| Claude-5 Sonnet (Vision) | 108 | 102 |
| Manus | 118 | 100 |
| Perplexity | 97 | 97 |
| DeepSeek V4 Pro | 108 | 88 |
| Bing Copilot | 94 | 83 |
| Mistral Large 3 | 88 | 74 |
| Kimi K3 | N/A | 130 |
| GLM 5.2 | N/A | 123 |
| Muse Glimmer | N/A | 123 |
| Muse Glimmer (Vision) | N/A | 110 |
Last updated, September 09, 2026
The latest AI IQ Test update has reshuffled the leaderboard again. Claude-5.1 Fable now holds the highest Mensa Norway score at 151, while GPT-5.6 TERRA Ultra (Vision) remains the leader on the offline test at 136.
The public Mensa leaderboard has also expanded sharply at the top. Thirteen models now score 140 or higher, compared with eight in the previous update. Claude-5.1 Fable leads at 151, followed by several models at 145, while Grok 4.6 High and Gemini 3.1 Pro reach 144.
Claude 5 Opus has improved enough that the previous “misses the genius tier” framing no longer applies. Standard Claude-5 Opus now scores 141 on Mensa Norway and 129 offline, while Claude-5 Opus MAX records 141 Mensa and 128 offline.
The offline leaderboard remains considerably more conservative. GPT-5.6 TERRA Ultra (Vision) leads at 136, GPT-5.6 TERRA Ultra follows at 133, and GPT-5.6 SOL Ultra (Vision) scores 132. No model has yet reached 140 on the offline test.
That contrast is central to understanding any AI IQ Test result. Mensa Norway shows how models perform on a familiar public test, while the unpublished offline set is intended to reduce the influence of test-data exposure. Rather than declaring one model universally “smarter,” this article examines what both scores reveal, why they sometimes disagree, and how much confidence we should place in an AI IQ leaderboard.
This essay is an attempt to sort the signal from the noise in our AI IQ Test era. I’ll walk through what human IQ actually measures, how researchers retrofit those tests for silicon brains to mimic an AI IQ level, what the AI IQ Test leaderboard really shows in mid-2026, and, just as important, what it hides behind the AI IQ Test façade. Along the way I’ll sprinkle in a few war stories from the lab, some philosophical detours on what is AI IQ, and a plea to keep our humility intact while machines crank out puzzle solutions at superhuman speed in these AI IQ 2026 contexts.
Table of Contents
1. A Brief History of Chasing the AI IQ Test Number
Psychologists have been quantifying intelligence for over a century, ever since Alfred Binet’s school-placement experiments evolved into the modern Intelligence Quotient and, much later, inspired today’s AI IQ Test comparisons. Set the population mean at 100 and the standard deviation at 15, and you get a tidy bell curve that places scores into recognizable ranges.
The number gained cultural influence because it was simple, standardized, and, in some contexts, predictive of outcomes such as academic performance. That simplicity also explains why researchers and readers are drawn to AI IQ scores, even though applying a human intelligence scale to a language model involves serious limitations.
Computers joined the AI IQ Test race only recently. Early neural networks in the 1990s would have performed near chance on many IQ-style tasks. Even GPT-3, which impressed users with fluent writing in 2020, reportedly landed around the high 70s when researchers informally tested it on Raven-style matrices. At the time, that placed it closer to a capable school student than a machine approaching the human genius range.
The latest leaderboard looks radically different.
On the leak-resistant offline test, GPT-5.6 TERRA Ultra (Vision) remains in first place with an IQ score of 136. GPT-5.6 TERRA Ultra follows at 133, while GPT-5.6 SOL Ultra (Vision) scores 132.
A group of models follows at 130: Claude-5.1 Fable, Claude-5.1 Fable (Vision), GPT-5.6 SOL Ultra, GPT-5.6 LUNA Max (Vision), and the offline-only Kimi K3 result. Claude-5 Opus sits just behind them at 129, while Grok 4.6 High, Claude-5 Opus MAX, and Gemini 3.7 Flash score 128.
The public Mensa Norway leaderboard is much more dramatic. Thirteen models now cross the chart’s 140-point genius threshold:
- Claude-5.1 Fable — 151
- Claude-5.1 Fable (Vision) — 145
- GPT-5.6 SOL Ultra — 145
- GPT-5.6 LUNA Max (Vision) — 145
- GPT-5.6 LUNA Max — 145
- Grok 4.6 High — 144
- Gemini 3.1 Pro — 144
- Gemini 3.7 Flash — 143
- GPT-5.6 TERRA Ultra — 142
- Qwen 3.8 Max — 142
- GPT-5.6 TERRA Ultra (Vision) — 141
- Claude-5 Opus — 141
- Claude-5 Opus MAX — 141
Claude-5.1 Fable is now the clear Mensa Norway leader at 151, a substantial jump above the previous 147-point ceiling. Yet its offline score is only 130, reinforcing the large gap that can appear between performance on a familiar public test and unpublished questions.
Claude 5 Opus has also improved enough that the previous “misses the genius tier” framing no longer applies. Claude-5 Opus now scores 141 on Mensa Norway and 129 offline, while Claude-5 Opus MAX scores 141 on Mensa Norway and 128 offline. Both therefore cross the public 140-point line, although neither challenges the leaders of the offline test.
The contrast between the two tests is exactly why this AI IQ Test leaderboard displays both scores side by side. The public Mensa test shows performance on a familiar and widely available question set, while the offline test is designed to reduce the influence of training-data exposure.
The results can differ substantially. A model may dominate Mensa Norway yet fall several places on the offline board, while another model may perform better on unpublished questions than on the public test.
Before we hail these systems as silicon Einsteins, we therefore need to separate two related but distinct questions: what IQ captures in human beings, and how faithfully an AI IQ Test measures anything comparable inside a language model.
The contrast between the two tests is exactly why this AI IQ Test leaderboard displays both scores side by side. The public Mensa test shows performance on a familiar and widely available question set, while the offline test is designed to reduce the influence of training-data exposure.
The results can differ substantially. A model may dominate Mensa Norway yet fall several places on the offline board, while another model may perform better on unpublished questions than on the public test.
Before we hail these systems as silicon Einsteins, we therefore need to separate two related but distinct questions: what IQ captures in human beings, and how faithfully an AI IQ Test measures anything comparable inside a language model.
2. What IQ Means for Flesh and Blood Thinkers in an AI IQ Test World
Ask five psychologists for an IQ definition and you’ll get six footnotes, but the consensus boils down to general cognitive ability — the fabled g factor central to any AI IQ Test. Classic batteries such as the Wechsler Adult Intelligence Scale (WAIS-IV) spread that umbrella over four pillars essential to AI IQ tests:
- Verbal Comprehension – vocabulary depth, analogies, general knowledge tested by AI IQ tests.
- Perceptual Reasoning – spatial puzzles, pattern completion, visual-spatial logic critical for an AI IQ Test.
- Working Memory – digit span, arithmetic under pressure, reflected in AI IQ Test results.
- Processing Speed – how fast you can chew through routine symbol matching, akin to google ai iq speed tests.
Scores are normed on massive, demographically balanced samples — tens of thousands of volunteers spanning ages, education levels, and cultures — much like IQ tests for humans. Every decade or so the publishers reform because societies slowly get better at test-taking (the famed Flynn effect), which parallels AI IQ 2026 recalibrations.
A critical property of a human IQ Test is that it stretches the brain in several directions at once, setting a standard for an AI IQ Test to follow. Try Raven’s matrices and you’ll feel the clock tick, your occipital cortex juggling shapes while your prefrontal cortex tracks rule candidates — similar to AI IQ Test prompts.
Switch to verbal analogies and you recruit entirely different neural circuits plus a lifetime of reading, paralleling what is AI IQ in transformers. The composite score therefore whispers something about how flexibly you think across domains, not just within one — a benchmark that AI IQ Test aims to emulate in AI IQ level metrics.
3. Translating the Exam for LLMs in AI IQ Test Context

Large language models don’t have retinas, fingers, or stress hormones, yet they face AI IQ Test challenges. They inhabit a leisurely universe where a thirty-second time limit is irrelevant and long-term memory can be simulated by a trillion training tokens. To shoehorn them into a human IQ framework — i.e. an AI IQ Test — researchers resort to verbalized versions of otherwise visual puzzles.
Take the Mensa Norway matrix: a 3×3 grid of abstract shapes with one square missing. For an LLM the image becomes a textual description — “Top row: a black arrow pointing right, then two arrows stacked, then a question mark” — followed by eight candidate answers spelled out in similar prose. The model picks the letter of the best match, illustrating AI IQ Test response methods. The conversion step is already a minefield for any AI IQ Test prompt engineer: which words you choose, how you order details, whether you mention color first or orientation first — all can bump accuracy by several percentage points in an AI IQ Test scenario.
Prompt engineering turns into prompt alchemy, essential for accurate AI IQ tests and google AI IQ benchmarks. And since the questions (at least in the public Mensa set) float freely on the open web, a model may have ingested them verbatim during pretraining — essentially taking the AI IQ Test with an answer key taped under the desk.
To compensate, platforms such as TrackingAI craft offline variants — fresh puzzles never published online, served from air-gapped servers for robust AI IQ Test validation. You can see the same pattern in the current board. Claude-5.1 Fable reaches 151 on Mensa Norway but only 130 offline, a 21-point gap. GPT-5.6 SOL Ultra scores 145 on Mensa and 130 offline. The relationship can also reverse: GPT-5.6 SOL Ultra (Vision) scores 124 on Mensa Norway but 132 offline. These gaps are a useful reminder that the two tests measure performance under different question sets and conditions, and that a spectacular public-test score does not automatically translate into the strongest offline result.
Still, even the leak-proof scores hover well above most humans, proving that AI IQ Test performance is no mere parlor trick.
4. Anatomy of an AI IQ Test Score — Why Do Language Models Ace What Used to Stump Them?

- Scale, Scale, Scale – Newer frontier models train on orders of magnitude more tokens than their predecessors, with an expanded context window that lets them hold an entire puzzle conversation in “working memory,” boosting AI IQ Test performance.
- Mixed Modality Embeddings – Even if inference is text-only, many models pretrain on image–text pairs, seeding a latent visual faculty that later helps decode verbalized diagrams in AI IQ Test scenarios.
- Self-Consistency Sampling – Instead of answering once, the model rolls the dice dozens of times, then votes on its own outputs, improving AI IQ Test consistency.
- Chain-of-Thought Fine-Tuning – Researchers now encourage models to “show their work.” Paradoxically, forcing an LLM to spell out step-by-step logic improves the final answer in AI IQ Test settings and exposes faulty jumps we can later debug.
Those tricks embody an R&D arms race in AI IQ Test innovation. Each new technique buys a handful of points; stack enough and you vault past the human median in the next AI IQ Test.
5. Cracks in the AI IQ Test Mirror
Yet IQ inflation has its dark corners in AI IQ Test results. Here are a few the hype cycle politely sidesteps:
- Prompt Sensitivity – I once swapped a single adjective — “slanted” for “angled” — in a matrix prompt and watched a model’s answer flip from correct to wrong in an AI IQ Test run. Humans are more robust to such noise in AI IQ tests.
- Metamemory vs. Understanding – LLMs sometimes describe the pattern better than they apply it. They can chatter about symmetry yet miss that the missing tile must be blank, not striped, in the AI IQ Test logic.
- One-Shot Brilliance, Multi-Step Fragility – On Google’s Humanity’s Last Exam, which strings several reasoning hops together, top models still limp well below human performance even in advanced AI IQ Test formats. Humans soar over 70%.
- No Stakes – A machine never gets test anxiety, hunger pangs, or sweaty palms. Remove those stressors from humans and their scores jump too, narrowing the ai iq level gap.
While IQ tests are a fascinating proxy, benchmarks like Humanity’s Last Exam (HLE) provide a deeper look at complex reasoning. For a full breakdown of Grok 4’s record-setting performance on that test.
6. Alternative Yardsticks Beyond a Single AI IQ Test
Because of those blind spots, the community is frantically building broader evaluation suites beyond a simple AI IQ Test metric:
- ARC AGI: 400 hand-crafted abstraction puzzles. Frontier reasoning models still nail only a small fraction of them; the average 10-year-old human hits 60%, reminding us that ai iq tests vary widely.
- MATH 500 & AIME: Competition-level algebra and geometry, where narrow fine-tuning pays off in AI IQ Test contexts far more than raw scale does.
- SWE-Bench & GitHub Bugs: Can a model patch real-world code? Scores linger around 50%, enough to excite CTOs but still miles from professional reliability in ai iq ranking scenarios.
- Ethical Twins & Value Alignment: Prototype tests ask models to rank moral dilemmas. Results swing wildly with prompt phrasing, indicating shaky meta-ethics in AI IQ Test influence.
7. Humans vs. Transformers: Apples and AI IQ Test Oranges
- Architecture – Our brains combine spike-timed neurons, chemistry, and plasticity honed by evolution. Transformers juggle linear algebra on GPU wafers in AI IQ Test simulations. They may converge on similar outputs, yet the journey there differs profoundly.
- Embodiment – A toddler learns “gravity” by dropping cereal on the floor; a model knows it only through textual snippets, a gap AI IQ Test can’t bridge. One has muscle memory and scraped knees; the other compresses patterns in high-dimensional space.
- Energy Budget – The human brain consumes about 20 W. Training a frontier LLM devours megawatt-hours. Efficiency is its own metric of intelligence — or at least survivability on a warming planet, which an AI IQ Test doesn’t measure.
Because of those chasms, IQ parity doesn’t imply cognitive parity in an AI IQ Test sense. If anything, it underscores how narrow tests can be hacked by alien architectures in AI IQ tests.
8. The Practical Outlook for an AI IQ Test–Driven World
I field this question weekly from product teams deciding whether to integrate the latest model into their AI IQ Test pipelines. My answer is a cautious yes, but:
- Yes, because higher puzzle competence often translates to crisper coding assistance, tighter logical reasoning, and fewer embarrassing math slips in your customer chatbot — real improvements in AI IQ level applications.
- But, because any decision with stakes — medical, financial, legal — still demands a human in the loop until we have evaluations that capture nuance, context, and moral reasoning beyond the AI IQ Test.
I like to frame IQ as “potential bandwidth.” It tells you the maximum data rate the channel can handle, not whether the message is truthful or safe. Your job is to wrap that channel in fail-safes, audits, and domain knowledge as part of your AI IQ Test deployment.
9. Where Testing Goes Next in the AI IQ Test Cycle
- Standardized Prompts – The community is coalescing around fixed, open-sourced prompt suites to curb cherry-picking in AI IQ Test creation.
- Multi-Modal Exams – Future tests will mix text, images, audio, maybe even robotics simulators, nudging models closer to the sensory buffet humans enjoy in AI IQ Test environments.
- Longitudinal Evaluations – Instead of a single snapshot, platforms like TrackingAI plan to probe the same model monthly, watching for “conceptual drift” as it undergoes post-deployment fine-tunes, ensuring AI IQ Test consistency.
- Alignment Leaderboards – Imagine a public scoreboard where models compete not for raw IQ but for harm reduction and truthfulness scores, far beyond a simple AI IQ Test. We’re early, but prototypes exist.
If we succeed, IQ will become just one cell in a sprawling spreadsheet — useful, but no longer in the limelight of AI IQ Test discourse.
10. Conclusion: Humility in the AI IQ Test Age of 151
Intelligence, in the richest sense, isn’t a single axis. It’s the braid of curiosity, empathy, street smarts, moral courage, creative spark, and, yes, raw reasoning speed, which an AI IQ Test only approximates. An LLM that slots puzzle pieces faster than 98% of humans has certainly achieved something historic in highest AI IQ history. But in my classroom, when that same student also helps a peer, questions an assumption, and owns a mistake — that’s when I nod and think, “There’s genius” beyond the AI IQ Test.
As AI marches up the IQ ladder, we should applaud the craftsmanship and absorb the lessons — then widen the lens to include AI IQ ranking and google AI IQ implications. Ask not just how bright the circuits glow, but where that light is pointed and whose face it illuminates. The future will be shaped by that broader definition of intelligence, one the AI IQ Test alone cannot capture.
Until then, enjoy the leaderboard. Just keep a saltshaker handy for those glittering numbers. They are map references, not the landscape itself.
Azmat — Founder of BinaryVerse AI | Tech Explorer and Observer of the Machine Mind Revolution
For questions or feedback, feel free to contact us or explore our About Us page
1) What is the average human IQ and where does the genius range start?
The average human IQ is standardized at 100. On this leaderboard, a score of 140 or higher is marked as the genius range. Thirteen models currently cross that threshold on Mensa Norway, but no model has reached 140 on the offline test.
2) Which AI has the highest IQ right now?
On the public Mensa Norway test, Claude-5.1 Fable currently has the highest score at 151. On the offline test, GPT-5.6 TERRA Ultra (Vision) leads with 136. Because the two tests use different question sets and testing conditions, there is no single uncontested measure of the highest AI IQ.
3) Why do Mensa and offline AI IQ scores differ
Mensa Norway uses publicly available questions that may have appeared in model training data. This can make the test partly measure familiarity as well as reasoning. The offline test uses unpublished questions and standardized testing conditions to reduce that risk.
Differences can also result from vision capabilities, prompt formatting, model tuning, and how well a particular model handles each test’s question style. That is why the leaderboard shows both scores rather than combining them into one number.
4) Which 13 models crossed the 140-point genius bar?
The 13 models scoring at least 140 on Mensa Norway are:
Claude-5.1 Fable — 151
Claude-5.1 Fable (Vision) — 145
GPT-5.6 SOL Ultra — 145
GPT-5.6 LUNA Max (Vision) — 145
GPT-5.6 LUNA Max — 145
Grok 4.6 High — 144
Gemini 3.1 Pro — 144
Gemini 3.7 Flash — 143
GPT-5.6 TERRA Ultra — 142
Qwen 3.8 Max — 142
GPT-5.6 TERRA Ultra (Vision) — 141
Claude-5 Opus — 141
Claude-5 Opus MAX — 141
None of these models has crossed 140 on the offline test.
5) How does Claude 5 Opus perform in the latest AI IQ Test?
Claude-5 Opus now scores 141 on Mensa Norway and 129 on the offline test. Claude-5 Opus MAX also reaches 141 on Mensa Norway, with an offline score of 128.
Both versions therefore cross the public test’s 140-point genius threshold. Their offline results remain competitive but sit below GPT-5.6 TERRA Ultra (Vision), which leads the offline leaderboard at 136.
6) Which models currently lead the offline AI IQ Test?
GPT-5.6 TERRA Ultra (Vision) leads the offline leaderboard with 136. It is followed by GPT-5.6 TERRA Ultra at 133 and GPT-5.6 SOL Ultra (Vision) at 132.
Five models currently sit at 130: Claude-5.1 Fable, Claude-5.1 Fable (Vision), GPT-5.6 SOL Ultra, GPT-5.6 LUNA Max (Vision), and Kimi K3. Kimi K3 currently has an offline result only, with no Mensa Norway score shown in the latest snapshot.
7) Are AI IQ scores a good predictor of coding or math performance?
AI IQ scores correlate with performance on abstract puzzles and pattern reasoning. That often helps with tasks like stepwise logic, careful arithmetic, and short algorithm sketches. However, production coding involves long-context planning, tool use, tests, and security concerns that IQ tests do not measure. If you care about software work, check task benchmarks such as competitive programming evaluations or code-fix leaderboards in addition to AI IQ rankings. Treat IQ as one input to model selection, not the only metric.
8) Can a single prompt or testing method change the score?
Yes. Prompt wording, answer format, and time or retry rules can shift accuracy. That is why our offline method uses standardized prompts and fixed scoring, and why we publish the test source next to each bar. If you see a large jump that only appears on one site or one day, look for differences in prompt phrasing or in how many chances the model received to answer. Stable methods produce stable rankings, which is the goal of showing Mensa and offline side by side.
9) Which source should I trust more, Mensa or offline?
Use them together. The Mensa Norway test shows how models handle a familiar public format, which can be influenced by training overlaps. Our offline set uses unpublished items and fixed prompts to reduce leakage, so it is a better view of generalization. If the two disagree, treat the offline number as the truer baseline and the Mensa number as a best case. That is why we show both bars and keep the chart date stamped.
10) How do you verify scores before adding them to the leaderboard?
Scores should only be added when they come from documented public testing or can be reproduced under consistent conditions. Mensa results should use the same question pool, prompt structure, answer format, and scoring method. Offline results should use unpublished questions, fixed prompts, and controlled scoring.
Models without a verified Mensa Norway score should remain marked as N/A rather than receiving an estimated value. In the current table, the models without a Mensa Norway result are Kimi K3, GLM 5.2, Muse Glimmer, and Muse Glimmer (Vision).

2 thoughts on “AI IQ Test September 2026: Claude 5.1 Fable Hits 151, 13 Models Cross 140”
Comments are closed.