AI IQ Test August 2026: 8 Models Break 140, but Claude 5 Opus Misses the Genius Tier

Last updated, August 1, 2026

ModelMensa Norway ScoreOffline Test Score
GPT 5.6 TERRA Ultra (Vision)141136
GPT 5.6 TERRA Ultra142134
GPT 5.6 SOL Ultra (Vision)124132
GPT 5.6 SOL Ultra144131
Grok 4.5 High147130
GPT 5.6 LUNA Max (Vision)145130
Claude-5 Opus MAX133130
Claude-5 Opus127130
Kimi-K3-VisionN/A130
Kimi K3N/A130
Claude-5 Fable (Vision)N/A130
Claude-5 Fable142129
Gemini 3.5 Flash Preview137128
GPT 5.6 LUNA Max147125
GLM 5.2N/A123
Muse Spark ThinkingN/A123
Gemini 3.1 Pro Preview145122
Gemini 3.1 Pro Preview (Vision)133117
Kimi K2.6136117
MiMo V2.5 Pro127116
Qwen 3.7130113
Grok 4.5 High (Vision)131111
Claude-5 SonnetN/A110
Muse Spark Thinking (Vision)N/A110
Manus118100
Perplexity9797
Claude-5 Sonnet (Vision)9790
DeepSeek V4 Pro10888
Bing Copilot9483
Mistral Large 38773
Data source: TrackingAI weekly IQ tests. See https://www.trackingai.org/home.

Last updated, July 18, 2026

The latest AI IQ Test results have produced the leaderboard’s biggest “genius tier” yet. Eight models now score 140 or higher on the public Mensa Norway test, led by Grok 4.5 High and GPT-5.6 LUNA Max, which share first place with scores of 147.

The other models crossing the 140-point line are GPT-5.6 LUNA Max (Vision) and Gemini 3.1 Pro Preview at 145, GPT-5.6 SOL Ultra at 144, GPT-5.6 TERRA Ultra and Claude-5 Fable at 142, and GPT-5.6 TERRA Ultra (Vision) at 141. That makes this the first update in which eight models simultaneously sit inside the chart’s designated genius range.

The arrival of Claude 5 Opus, however, is less spectacular on Mensa Norway. Claude-5 Opus MAX scores 133, while the standard Claude-5 Opus reaches 127, leaving both comfortably below the 140 threshold. Calling the Opus family an outright failure would be misleading, though: both versions score 130 on the offline test, placing them inside a crowded group of strong performers.

The leak-resistant offline leaderboard remains more conservative. GPT-5.6 TERRA Ultra (Vision) leads with 136, followed by GPT-5.6 TERRA Ultra at 134, GPT-5.6 SOL Ultra (Vision) at 132, and GPT-5.6 SOL Ultra at 131. No model has yet crossed 140 offline.

That contrast is central to understanding any AI IQ Test result. Mensa Norway shows how models perform on a familiar public test, while the unpublished offline set is intended to reduce the influence of test-data exposure. Rather than declaring one model universally “smarter,” this article examines what both scores reveal, why they sometimes disagree, and how much confidence we should place in an AI IQ leaderboard.

This essay is an attempt to sort the signal from the noise in our AI IQ Test era. I’ll walk through what human IQ actually measures, how researchers retrofit those tests for silicon brains to mimic an AI IQ level, what the AI IQ Test leaderboard really shows in mid-2026, and, just as important, what it hides behind the AI IQ Test façade. Along the way I’ll sprinkle in a few war stories from the lab, some philosophical detours on what is AI IQ, and a plea to keep our humility intact while machines crank out puzzle solutions at superhuman speed in these AI IQ 2026 contexts.

1. A Brief History of Chasing the AI IQ Test Number

Psychologists have been quantifying intelligence for over a century, ever since Alfred Binet’s school-placement experiments evolved into the modern Intelligence Quotient and, much later, inspired today’s AI IQ Test comparisons. Set the population mean at 100 and the standard deviation at 15, and you get a tidy bell curve that places scores into recognizable ranges.

The number gained cultural influence because it was simple, standardized, and, in some contexts, predictive of outcomes such as academic performance. That simplicity also explains why researchers and readers are drawn to AI IQ scores, even though applying a human intelligence scale to a language model involves serious limitations.

Computers joined the AI IQ Test race only recently. Early neural networks in the 1990s would have performed near chance on many IQ-style tasks. Even GPT-3, which impressed users with fluent writing in 2020, reportedly landed around the high 70s when researchers informally tested it on Raven-style matrices. At the time, that placed it closer to a capable school student than a machine approaching the human genius range.

The latest leaderboard looks radically different.

On the leak-resistant offline test, GPT-5.6 TERRA Ultra (Vision) now holds first place with an IQ score of 136. It is followed by GPT-5.6 TERRA Ultra at 134, GPT-5.6 SOL Ultra (Vision) at 132, and GPT-5.6 SOL Ultra at 131.

A crowded group of models follows at 130. That group includes Grok 4.5 High, GPT-5.6 LUNA Max (Vision), Claude-5 Opus MAX, Claude-5 Opus, Kimi-K3-Vision, Kimi K3, and Claude-5 Fable (Vision). Claude-5 Fable sits just behind them with an offline score of 129, while Gemini 3.5 Flash Preview reaches 128.

The public Mensa Norway leaderboard tells a more dramatic story. Eight models now cross the chart’s 140-point genius threshold:

  • Grok 4.5 High — 147
  • GPT-5.6 LUNA Max — 147
  • GPT-5.6 LUNA Max (Vision) — 145
  • Gemini 3.1 Pro Preview — 145
  • GPT-5.6 SOL Ultra — 144
  • GPT-5.6 TERRA Ultra — 142
  • Claude-5 Fable — 142
  • GPT-5.6 TERRA Ultra (Vision) — 141

The most closely watched new addition is the Claude 5 Opus family, but its Mensa Norway debut is less impressive than the surrounding hype might suggest. Claude-5 Opus MAX scores 133, while the standard Claude-5 Opus reaches only 127. Both fall well short of the 140 genius line.

Their offline performance is more competitive. Claude-5 Opus MAX and Claude-5 Opus both score 130, tying several of the strongest models on the leak-resistant test. The fairest interpretation is therefore not that Opus 5 completely failed, but that it failed to make the expected impact on the public Mensa leaderboard while remaining a serious offline contender.

The contrast between the two tests is exactly why this AI IQ Test leaderboard displays both scores side by side. The public Mensa test shows performance on a familiar and widely available question set, while the offline test is designed to reduce the influence of training-data exposure.

The results can differ substantially. A model may dominate Mensa Norway yet fall several places on the offline board, while another model may perform better on unpublished questions than on the public test.

Before we hail these systems as silicon Einsteins, we therefore need to separate two related but distinct questions: what IQ captures in human beings, and how faithfully an AI IQ Test measures anything comparable inside a language model.

2. What IQ Means for Flesh and Blood Thinkers in an AI IQ Test World

Ask five psychologists for an IQ definition and you’ll get six footnotes, but the consensus boils down to general cognitive ability — the fabled g factor central to any AI IQ Test. Classic batteries such as the Wechsler Adult Intelligence Scale (WAIS-IV) spread that umbrella over four pillars essential to AI IQ tests:

  • Verbal Comprehension – vocabulary depth, analogies, general knowledge tested by AI IQ tests.
  • Perceptual Reasoning – spatial puzzles, pattern completion, visual-spatial logic critical for an AI IQ Test.
  • Working Memory – digit span, arithmetic under pressure, reflected in AI IQ Test results.
  • Processing Speed – how fast you can chew through routine symbol matching, akin to google ai iq speed tests.

Scores are normed on massive, demographically balanced samples — tens of thousands of volunteers spanning ages, education levels, and cultures — much like IQ tests for humans. Every decade or so the publishers reform because societies slowly get better at test-taking (the famed Flynn effect), which parallels AI IQ 2026 recalibrations.

A critical property of a human IQ Test is that it stretches the brain in several directions at once, setting a standard for an AI IQ Test to follow. Try Raven’s matrices and you’ll feel the clock tick, your occipital cortex juggling shapes while your prefrontal cortex tracks rule candidates — similar to AI IQ Test prompts.

Switch to verbal analogies and you recruit entirely different neural circuits plus a lifetime of reading, paralleling what is AI IQ in transformers. The composite score therefore whispers something about how flexibly you think across domains, not just within one — a benchmark that AI IQ Test aims to emulate in AI IQ level metrics.

3. Translating the Exam for LLMs in AI IQ Test Context

AI IQ Test leaderboard cover image showing GPT-5.6 Ultra leading the offline benchmark
AI IQ Test leaderboard cover image showing GPT-5.6 Ultra leading the offline benchmark

Large language models don’t have retinas, fingers, or stress hormones, yet they face AI IQ Test challenges. They inhabit a leisurely universe where a thirty-second time limit is irrelevant and long-term memory can be simulated by a trillion training tokens. To shoehorn them into a human IQ framework — i.e. an AI IQ Test — researchers resort to verbalized versions of otherwise visual puzzles.

Take the Mensa Norway matrix: a 3×3 grid of abstract shapes with one square missing. For an LLM the image becomes a textual description — “Top row: a black arrow pointing right, then two arrows stacked, then a question mark” — followed by eight candidate answers spelled out in similar prose. The model picks the letter of the best match, illustrating AI IQ Test response methods. The conversion step is already a minefield for any AI IQ Test prompt engineer: which words you choose, how you order details, whether you mention color first or orientation first — all can bump accuracy by several percentage points in an AI IQ Test scenario.

Prompt engineering turns into prompt alchemy, essential for accurate AI IQ tests and google AI IQ benchmarks. And since the questions (at least in the public Mensa set) float freely on the open web, a model may have ingested them verbatim during pretraining — essentially taking the AI IQ Test with an answer key taped under the desk.

To compensate, platforms such as TrackingAI craft offline variants — fresh puzzles never published online, served from air-gapped servers for robust AI IQ Test validation. You can see the same pattern in the current board: Grok-4.20 Expert Mode posts 144 on Mensa but only 122 offline, while Claude-4.8 Opus posts 141 on Mensa but 117 offline. The most extreme case in the current dataset is Claude-4.8 Opus MAX (Vision), which scores a solid 123 offline yet crashes to 67 on Mensa — the reverse pattern from what we usually see, and a good reminder that these gaps can run in either direction depending on how a model was tuned.

Still, even the leak-proof scores hover well above most humans, proving that AI IQ Test performance is no mere parlor trick.

4. Anatomy of an AI IQ Test Score — Why Do Language Models Ace What Used to Stump Them?

 Layered infographic explaining the four factors behind a high AI IQ Test score in language models
Layered infographic explaining the four factors behind a high AI IQ Test score in language models
  1. Scale, Scale, Scale – Newer frontier models train on orders of magnitude more tokens than their predecessors, with an expanded context window that lets them hold an entire puzzle conversation in “working memory,” boosting AI IQ Test performance.
  2. Mixed Modality Embeddings – Even if inference is text-only, many models pretrain on image–text pairs, seeding a latent visual faculty that later helps decode verbalized diagrams in AI IQ Test scenarios.
  3. Self-Consistency Sampling – Instead of answering once, the model rolls the dice dozens of times, then votes on its own outputs, improving AI IQ Test consistency.
  4. Chain-of-Thought Fine-Tuning – Researchers now encourage models to “show their work.” Paradoxically, forcing an LLM to spell out step-by-step logic improves the final answer in AI IQ Test settings and exposes faulty jumps we can later debug.

Those tricks embody an R&D arms race in AI IQ Test innovation. Each new technique buys a handful of points; stack enough and you vault past the human median in the next AI IQ Test.

5. Cracks in the AI IQ Test Mirror

Yet IQ inflation has its dark corners in AI IQ Test results. Here are a few the hype cycle politely sidesteps:

  • Prompt Sensitivity – I once swapped a single adjective — “slanted” for “angled” — in a matrix prompt and watched a model’s answer flip from correct to wrong in an AI IQ Test run. Humans are more robust to such noise in AI IQ tests.
  • Metamemory vs. Understanding – LLMs sometimes describe the pattern better than they apply it. They can chatter about symmetry yet miss that the missing tile must be blank, not striped, in the AI IQ Test logic.
  • One-Shot Brilliance, Multi-Step Fragility – On Google’s Humanity’s Last Exam, which strings several reasoning hops together, top models still limp well below human performance even in advanced AI IQ Test formats. Humans soar over 70%.
  • No Stakes – A machine never gets test anxiety, hunger pangs, or sweaty palms. Remove those stressors from humans and their scores jump too, narrowing the ai iq level gap.

While IQ tests are a fascinating proxy, benchmarks like Humanity’s Last Exam (HLE) provide a deeper look at complex reasoning. For a full breakdown of Grok 4’s record-setting performance on that test.

6. Alternative Yardsticks Beyond a Single AI IQ Test

Because of those blind spots, the community is frantically building broader evaluation suites beyond a simple AI IQ Test metric:

  • ARC AGI: 400 hand-crafted abstraction puzzles. Frontier reasoning models still nail only a small fraction of them; the average 10-year-old human hits 60%, reminding us that ai iq tests vary widely.
  • MATH 500 & AIME: Competition-level algebra and geometry, where narrow fine-tuning pays off in AI IQ Test contexts far more than raw scale does.
  • SWE-Bench & GitHub Bugs: Can a model patch real-world code? Scores linger around 50%, enough to excite CTOs but still miles from professional reliability in ai iq ranking scenarios.
  • Ethical Twins & Value Alignment: Prototype tests ask models to rank moral dilemmas. Results swing wildly with prompt phrasing, indicating shaky meta-ethics in AI IQ Test influence.

7. Humans vs. Transformers: Apples and AI IQ Test Oranges

  • Architecture – Our brains combine spike-timed neurons, chemistry, and plasticity honed by evolution. Transformers juggle linear algebra on GPU wafers in AI IQ Test simulations. They may converge on similar outputs, yet the journey there differs profoundly.
  • Embodiment – A toddler learns “gravity” by dropping cereal on the floor; a model knows it only through textual snippets, a gap AI IQ Test can’t bridge. One has muscle memory and scraped knees; the other compresses patterns in high-dimensional space.
  • Energy Budget – The human brain consumes about 20 W. Training a frontier LLM devours megawatt-hours. Efficiency is its own metric of intelligence — or at least survivability on a warming planet, which an AI IQ Test doesn’t measure.

Because of those chasms, IQ parity doesn’t imply cognitive parity in an AI IQ Test sense. If anything, it underscores how narrow tests can be hacked by alien architectures in AI IQ tests.

8. The Practical Outlook for an AI IQ Test–Driven World

I field this question weekly from product teams deciding whether to integrate the latest model into their AI IQ Test pipelines. My answer is a cautious yes, but:

  • Yes, because higher puzzle competence often translates to crisper coding assistance, tighter logical reasoning, and fewer embarrassing math slips in your customer chatbot — real improvements in AI IQ level applications.
  • But, because any decision with stakes — medical, financial, legal — still demands a human in the loop until we have evaluations that capture nuance, context, and moral reasoning beyond the AI IQ Test.

I like to frame IQ as “potential bandwidth.” It tells you the maximum data rate the channel can handle, not whether the message is truthful or safe. Your job is to wrap that channel in fail-safes, audits, and domain knowledge as part of your AI IQ Test deployment.

9. Where Testing Goes Next in the AI IQ Test Cycle

  1. Standardized Prompts – The community is coalescing around fixed, open-sourced prompt suites to curb cherry-picking in AI IQ Test creation.
  2. Multi-Modal Exams – Future tests will mix text, images, audio, maybe even robotics simulators, nudging models closer to the sensory buffet humans enjoy in AI IQ Test environments.
  3. Longitudinal Evaluations – Instead of a single snapshot, platforms like TrackingAI plan to probe the same model monthly, watching for “conceptual drift” as it undergoes post-deployment fine-tunes, ensuring AI IQ Test consistency.
  4. Alignment Leaderboards – Imagine a public scoreboard where models compete not for raw IQ but for harm reduction and truthfulness scores, far beyond a simple AI IQ Test. We’re early, but prototypes exist.

If we succeed, IQ will become just one cell in a sprawling spreadsheet — useful, but no longer in the limelight of AI IQ Test discourse.

10. Conclusion: Humility in the AI IQ Test Age of 136

Intelligence, in the richest sense, isn’t a single axis. It’s the braid of curiosity, empathy, street smarts, moral courage, creative spark, and, yes, raw reasoning speed, which an AI IQ Test only approximates. An LLM that slots puzzle pieces faster than 98% of humans has certainly achieved something historic in highest AI IQ history. But in my classroom, when that same student also helps a peer, questions an assumption, and owns a mistake — that’s when I nod and think, “There’s genius” beyond the AI IQ Test.

As AI marches up the IQ ladder, we should applaud the craftsmanship and absorb the lessons — then widen the lens to include AI IQ ranking and google AI IQ implications. Ask not just how bright the circuits glow, but where that light is pointed and whose face it illuminates. The future will be shaped by that broader definition of intelligence, one the AI IQ Test alone cannot capture.

Until then, enjoy the leaderboard. Just keep a saltshaker handy for those glittering numbers. They are map references, not the landscape itself.

Azmat — Founder of BinaryVerse AI | Tech Explorer and Observer of the Machine Mind Revolution
For questions or feedback, feel free to contact us or explore our About Us page

AI IQ Testing:
Involves researchers adapting traditional human intelligence tests, such as matrix tests, and applying them to artificial intelligence models, particularly large language models (LLMs). This practice has gained prominence as recent AI models have achieved high scores on these tests.
Large Language Models (LLMs):
A type of artificial intelligence model that processes and generates text. Recent advancements in LLMs have led to significant increases in their performance on adapted IQ tests.
Intelligence Quotient (IQ):
A numerical score intended to quantify intelligence. Traditionally, human IQ scores are set with a population mean of 100 and a standard deviation of 15, following a bell curve. AI models are now being given “AI IQ” scores based on their performance on adapted tests.
g factor (general cognitive ability):
A consensus view in psychology holds that human IQ measures this underlying general cognitive ability. Classic human IQ tests like the WAIS IV aim to assess this across various domains.
Wechsler Adult Intelligence Scale (WAIS IV):
A classic battery of human IQ tests that assesses general cognitive ability across four main pillars: Verbal Comprehension, Perceptual Reasoning, Working Memory, and Processing Speed.

1) What is the average human IQ and where does the genius range start?

The average human IQ is standardized at 100. On this leaderboard, a score of 140 or higher is marked as the genius range. Eight models currently cross that threshold on Mensa Norway, but no model has reached 140 on the offline test.

2) Which AI has the highest IQ right now?

On the public Mensa Norway test, Grok 4.5 High and GPT-5.6 LUNA Max share the highest score at 147. On the offline test, GPT-5.6 TERRA Ultra (Vision) leads with 136. Because the two tests use different question sets and testing conditions, there is no single uncontested highest AI IQ.

3) Why do Mensa and offline AI IQ scores differ

Mensa Norway uses publicly available questions that may have appeared in model training data. This can make the test partly measure familiarity as well as reasoning. The offline test uses unpublished questions and standardized testing conditions to reduce that risk.
Differences can also result from vision capabilities, prompt formatting, model tuning, and how well a particular model handles each test’s question style. That is why the leaderboard shows both scores rather than combining them into one number.

4) Which eight models crossed the 140-point genius bar?

The eight models scoring at least 140 on Mensa Norway are:
Grok 4.5 High — 147
GPT-5.6 LUNA Max — 147
GPT-5.6 LUNA Max (Vision) — 145
Gemini 3.1 Pro Preview — 145
GPT-5.6 SOL Ultra — 144
GPT-5.6 TERRA Ultra — 142
Claude-5 Fable — 142
GPT-5.6 TERRA Ultra (Vision) — 141
None of these models has crossed 140 on the offline test.

5) Did Claude 5 Opus fail to impress in the latest AI IQ Test?

Claude 5 Opus missed the headline-making Mensa genius tier. Claude-5 Opus MAX scored 133 on Mensa Norway, while the standard Claude-5 Opus scored 127.
Its offline performance was stronger: both versions scored 130, tying several other leading models. The fairest conclusion is that Claude 5 Opus underperformed expectations on the public Mensa test but remained competitive on the leak-resistant offline benchmark.

6) Which models currently lead the offline AI IQ Test?

GPT-5.6 TERRA Ultra (Vision) leads the offline leaderboard with 136. It is followed by GPT-5.6 TERRA Ultra at 134, GPT-5.6 SOL Ultra (Vision) at 132, and GPT-5.6 SOL Ultra at 131.
A large group then ties at 130, including Grok 4.5 High, GPT-5.6 LUNA Max (Vision), Claude-5 Opus MAX, Claude-5 Opus, Kimi-K3-Vision, Kimi K3, and Claude-5 Fable (Vision). Because so many models share 130, presenting a strict “Top 10” would require arbitrarily excluding one of the tied models.

7) Are AI IQ scores a good predictor of coding or math performance?

AI IQ scores correlate with performance on abstract puzzles and pattern reasoning. That often helps with tasks like stepwise logic, careful arithmetic, and short algorithm sketches. However, production coding involves long-context planning, tool use, tests, and security concerns that IQ tests do not measure. If you care about software work, check task benchmarks such as competitive programming evaluations or code-fix leaderboards in addition to AI IQ rankings. Treat IQ as one input to model selection, not the only metric.

8) Can a single prompt or testing method change the score?

Yes. Prompt wording, answer format, and time or retry rules can shift accuracy. That is why our offline method uses standardized prompts and fixed scoring, and why we publish the test source next to each bar. If you see a large jump that only appears on one site or one day, look for differences in prompt phrasing or in how many chances the model received to answer. Stable methods produce stable rankings, which is the goal of showing Mensa and offline side by side.

9) Which source should I trust more, Mensa or offline?

Use them together. The Mensa Norway test shows how models handle a familiar public format, which can be influenced by training overlaps. Our offline set uses unpublished items and fixed prompts to reduce leakage, so it is a better view of generalization. If the two disagree, treat the offline number as the truer baseline and the Mensa number as a best case. That is why we show both bars and keep the chart date stamped.

10) How do you verify scores before adding them to the leaderboard?

Scores should only be added when they come from documented public testing or can be reproduced under consistent conditions. Mensa results should use the same question pool, prompt structure, answer format, and scoring method. Offline results should use unpublished questions, fixed prompts, and controlled scoring.
Models without a verified Mensa Norway score should remain marked as N/A rather than receiving an estimated value. In the current table, these include Kimi-K3-Vision, Kimi K3, Claude-5 Fable (Vision), GLM 5.2, Muse Spark Thinking, Claude-5 Sonnet, and Muse Spark Thinking (Vision).

2 thoughts on “AI IQ Test August 2026: 8 Models Break 140, but Claude 5 Opus Misses the Genius Tier”

  1. Pingback: Denksplitter

Comments are closed.