AI vs. Logic Puzzles: Has AI Gotten Any Better?

In 2024, Carnegie Mellon University associate professor Xaq Pitkow’s claims about AI made headlines: “It tends to be worse than humans at questions that require more abstract thinking.” According to him, AI was great at recognizing patterns but did terribly at reasoning.1

But this was back when ChatGPT’s latest model was still GPT-4, Anthropic’s Claude AI was only starting to gain ground, and Google Gemini had just been newly released. Given the speed at which these AI models have been developing, we wanted to answer a simple question: have today’s AI models gotten better at solving logic puzzles?

What We Tested: AI Accuracy and Efficiency in Puzzles

Our team wanted to test whether newer AI models are capable of successfully solving different types of logic puzzles. These included challenges that require finding the right words, arranging numbers in the right way, calculating correct answers, or coloring the right squares.

Our team tested Claude 4.6, Gemini 3, ChatGPT 5.3, and Grok 4.1, asking each model to answer one daily challenge for each puzzle for a month (31 days). We looked at how efficient each AI model was for each type of puzzle. How long did it take them to arrive at the correct answer? Did they manage it on the first try (or at all)?

We logged the following metrics:

    • Average Solve Rate (ASR): The percentage of puzzles for which the model produced the complete and correct final solution grid.
        • ASR (1st): Success rate on the first attempt.
        • ASR (2nd): Success rate on the second attempt (when the first attempt failed).
        • ASR (total): Overall success rate (solved correctly on either the first or second try).
    • Average Solve Time (AST): How long it took for each model to return the correct answers. This only includes successful solves, not time spent on incorrect attempts.
    • Average Number of Tries (ANT): The average number of attempts needed to reach the correct answer on Wordle, where only 6 guesses are allowed by the game. (For the rest of the puzzles, the models were given two tries only.)

Jump to our Appendix to read more about our methodology, including detailed information about each puzzle, how the tests differed slightly depending on the task, and the limitations of the study.

What We Found: AI Still Struggles With Multi-Step Reasoning

While the AI models tested did relatively well on smaller, more constrained, and familiar puzzles (such as Wordle or The Mini crossword), they generally performed worse in harder puzzles or under stricter prompts. 

Even when they got the answers right for slightly more difficult challenges, they often couldn’t explain how they arrived at the solution. They also struggled with open-ended puzzles, such as Connections. We observed that the models often relied on guesswork and answer retrieval (from published online sources), and in several cases self-reported doing so, rather than showing a clear step-by-step path to the solution.

Grok had the highest overall solve rate. It relatively struggled with spatial puzzles, failing to correctly answer around half of the easy and slightly harder puzzles, but otherwise managed to solve at least 84% of the puzzles. This model also did the best at language and crossword puzzles. It returned the correct answer around 96% of the time and took an average of only 24 seconds to solve them. 

Claude was similarly consistent across categories and difficulties. It arrived at the correct solution for 4 puzzles within the allowed number of attempts (up to 6 for Wordle and up to 2 for Futoshiki, MathemaGrids, and Nonogrids). This model often succeeded on the second try, where other AI models gave up. 

Gemini and ChatGPT were middle-to-lower performers, with the latter frequently coming in last across multiple categories.

Notably, all models often made mistakes on the first try but got the correct answer once asked to try again. This shows that they can recover when given chances or guided, but they still miss obvious errors on their own and need human verification.

We arrived at the conclusions above by reviewing the trends present in each puzzle category.

Language Puzzles

Language-based puzzles ask users to find or connect random words based on logical links. All three puzzles here require basic knowledge of the English language, such as spellings, definitions, and context clues.

Grok was the only model to solve all Wordle puzzles correctly on the first try. However, it admitted that it doesn’t “think” about its solutions. Instead, it extracts answers from wikis, gaming sites, and archives.

Claude did the same. Like Grok, it achieved a perfect solve rate by looking up the answers, except that it took 3 tries on average. When asked about it, Claude described spacing out its guesses to resemble a human’s typical guessing pattern, rather than submitting the answer on the first try. It estimated that human players would need at least 3 guesses to arrive at the correct answer and paced its responses accordingly.

ChatGPT did the worst across the language category. It only solved a little over half of the Connections puzzles, 80.7% of the Daily Puzzles, and just a quarter of the Wordle puzzles (taking about 5 attempts on average to get to the right answer).

All in all, playing Wordle by the rules seemed to be a struggle for the four AI models we tested. Grok and Claude both reported pulling the correct answer from external sources rather than deriving it through gameplay, while Gemini and ChatGPT often responded with lists of multiple possible answers instead of giving a definite guess each time. So, even when they eventually reached the correct answer, they didn’t follow a methodical or logical pathway.

Our team noticed that AI models tended to violate clear primary constraints. In one instance, Claude submitted a 6-letter word for Wordle, even when the answer is always exactly 5-letters long. 

For Connections, the models would sometimes place the same word in two different categories, which is a clear violation of the game’s parameters. It was also common for them to group 3 words correctly, and then get the fourth word wrong.

On the other hand, all 4 models did relatively well at the Daily Puzzle, despite human categorization placing it under the “slightly harder” difficulty group. These narrative, riddle-style puzzles typically require pattern recognition, which might explain why the AI models handled them better than they did the stricter grid-based challenges.

Numerical Puzzles

Puzzles under this category require the user to arrange a limited set of numbers in a grid so that algebraic conditions are met. A player would need elementary knowledge of numbers and basic arithmetic. Some of these challenges don’t follow established mathematical rules (such as the order of operations). Instead, the puzzles specify their own rules and constraints.

The AI models did relatively well at solving the easy and medium-difficulty numerical puzzles on the first try. 3 models managed to get a perfect solve rate in a puzzle without needing to retry: Grok for Futoshiki and Claude and ChatGPT for MathemaGrids.

Claude performed the best. It only failed to get a 100% total ASR for CalcuDoku, scoring 96.77% instead (still the highest among the models). It also consistently got the highest or second-highest first-try ASR across all three games. This suggests that Claude is more adept at solving basic math than other models, though it wasn’t the best at word play.

On some occasions, however, the AI models committed more obvious constraint violations. For Futoshiki, they sometimes tried to overwrite pre-filled squares (clues), repeat numbers in the same row or column, or fail the basic “greater than” or “less than” constraints.

We observed similar violations in MathemaGrids and CalcuDoku, where they would mess up the order of operations, ignore grid conditions, or use the wrong mathematical operation altogether (such as adding instead of multiplying).

Gemini was the overall worst performer, managing to solve Futoshiki and CalcuDoku puzzles only 29% of the time.

Spatial Puzzles

Spatial puzzles require the user to color or mark the right squares in a grid so that the game’s conditions, of which there are typically several layers, are met. In 3-In-A-Row, for instance, each row and column needs to have an equal number of blue and white squares, and there can’t be three of the same color next to each other. As the puzzle’s difficulty increases, so do the constraints.

The AI models seemed to struggle most with spatial puzzles. Only Claude managed to get a 100% ASR in any of the puzzles and maintained an overall average of 92.5%. Notably, Gemini failed to correctly answer any Daily Tents puzzle and only managed to solve a few 3-in-A-Row puzzles on the second try.

The mistakes made by the AI models were a lot more glaring in this category, too. They often violated the most prominent rules or constraints of the puzzles and failed to recognize wrong answers unless prompted by a human.

For 3-In-A-Row, they repeatedly broke the only rule (not having three consecutive legends) and overwrote pre-filled cells. For Nonogrids, the models would often fill in extra blocks where none should exist, get block groupings wrong, and fail to leave the required gaps between groups.

The same was seen with Daily Tents. Although it was clearly indicated in the starting prompt that tents can’t be next to each other, the models kept placing adjacent (or even overlapping) tents. Their answers also had plenty of missing tents, extra tents, or incorrect tent placements.

3-In-A-Row, despite being categorized as an easy puzzle, got the third-lowest overall ASR across models. Claude scored 90.3% for this puzzle, but the total was brought down significantly by Gemini and ChatGPT3 (getting only 12.9% each).

Daily Tents turned out to be the second-most difficult puzzle we tested (33.9% overall ASR), surpassed only by The Crossword. 

It’s possible that the layered constraints made it hard for AI models to satisfy all conditions while maintaining a clear logical path. Unlike language, crossword, or numerical puzzles (where existing knowledge or established rules help), spatial puzzles may require deeper reasoning.

Crossword Puzzles

Crossword puzzles require the user to find the word that corresponds to the given clue and arrange it in a grid. Unlike language-based puzzles, crossword clues typically deal with random facts or definitions and aren’t limited by the English language.

This is the only category where we observed a clear difficulty gradient that followed our human categorization of the puzzles. The overall ASR was 89.5% for The Mini, 44.4% for The Midi, and only 29.8% for The Crossword.

We saw the same drop in individual ASRs for all models except Grok, which scored 6.5 percentage points higher in The Crossword (90.3%) than in The Midi (83.9%). However, the model again admitted to using published sources, rather than step-by-step reasoning, to extract its answers. This doesn’t explain why it failed to give the correct solution for all puzzles.

In fact, on 3 separate occasions, Grok took minutes of “thinking” only to return “no response” or “an error occurred.” We didn’t see this with any other model.

The difficulty in solving crossword puzzles was more pronounced with Claude and Gemini. After getting a total ASR of 90.3% and 83.9%, respectively, for The Mini, their scores dropped to 16.1% (Claude) and 6.5% (Gemini) for The Midi. Moreover, Gemini consistently failed to return any correct answers on the second try.

The most common mistake we saw was outright clue mismatch, where the answer didn’t fit or solve the given clue. In one puzzle, for example, Claude answered “KATE (Cate Blanchett?)” to the clue “One Battle After Another star.” The puzzle was looking for a 3-letter word: “LEO.”

And, as the puzzles got bigger, the models tended to either mix up answers for multiple items (switching words) or even provide answers that were completely outside of a puzzle’s given theme.

The models would also mistakenly fill black squares (where no letter should be at all) or return a puzzle with a misplaced black square. In general, errors became much more frequent as the crossword size increased.

What the Data Says: Understanding What AI Is and Isn’t

In 2025, OpenAI published a story about how GPT-5 helped mathematician Ernest Ryu answer a 40-year-old open problem. According to Ryu, the feeling was “surreal” and comparable to “working with a competent student who would come up with ideas, ask questions, and brainstorm.2

However, the results of our study paint a slightly different picture. While AI can be used to solve problems with clear patterns and parameters, it markedly can’t think or process the way a human brain does. It doesn’t even seem to interpret difficulty levels the way we do.

Moreover, the models all committed obvious constraint violations and failed to recognize mistakes in their own work unless prompted by a human. The implications of this go beyond logic puzzles.

Regular users who rely on AI for problem-solving or decision-making might not double-check or question AI’s answers. Especially given that we’ve seen AI modify parameters and violate explicit conditions to “solve” challenges, users could unknowingly be getting misinformed or inaccurate responses.

At least until 2023, GPT-4 couldn’t reliably do basic arithmetic, like addition and multiplication; and when asked to show its work, it tended to hallucinate explanations or return inaccuracies.3 Moreover, an Apple study found that, as of 2025, most models “give up” when tasked with difficult human puzzles.4

Because of these observations, experts posited that generative AI models can often calculate the “immediate best moves” when solving a puzzle. However, they can’t perform the bigger logical processing that’s usually needed in longer, more complex games.5

This is because generative AI models are predictive text engines. They’re trained on massive datasets that allow them to recognize patterns and use those to create new content. They aren’t secure calculators, nor can they process complex logical challenges without targeted training.

So, while AI can be useful assistance for various activities, it’s incapable of fully replacing independent human processing just yet.

Why It Matters: What Responsible AI Use Looks Like

How adults in the US view and use AI: 59% feel they can't control AI's presence in their lives, 55% want more control; 27% admit to using AI extensively, 79% of experts think most adults in the US do.

In a survey by Pew Research Center, 59% of adults in the US expressed that they don’t feel in control of whether AI is used in their lives, and nearly the same number (55%) say that they do want more control.6 But with the rapid adoption of AI in consumer-facing businesses, tech, and digital experiences, that “control” seems to just be getting harder to achieve.

For instance, 57% of users admitted to interacting with AI at least once daily, with nearly 30% using it several times a day. However, nearly 80% of AI experts believe that most adults in the US use models “almost constantly.”6

Given how prevalent AI has become, the first step to using it in a safer, more controlled way is learning its limitations. Understanding that AI won’t always give you the correct answer can help you avoid making critical mistakes at work, in school, or in your daily life.

Generative models being unable to independently “think” using human reasoning also has cybersecurity implications. You shouldn’t expect AI to operate the way a person would, especially in terms of privacy and security. Never treat AI chatbots and programs like a secure vault for sensitive information, and review the privacy policies or statements of any AI model you use. If possible, opt for a privacy-focused model with strong encryption protocols.

Summing Up: Staying Informed and Protected in the Age of AI

The results of our study show that, despite recent advances in AI models, the technology still isn’t as adept at abstract reasoning or human logical processing as many would expect. In many cases, the models needed to be told that their answers were wrong and asked to try again before producing the correct result. This raises questions about the reliability of AI for complex problem-solving or decision-making.

But the use of generative AI is something most people can’t avoid. Many businesses and workplaces have integrated AI into their systems and require employees to use it, while digital products have become increasingly integrated with this tech.

What’s important is to be aware of what AI can and can’t do and, in turn, what it could and couldn’t be trusted to do. Being more intentional and vigilant when using AI can help you maximize the technology while staying safe, both in professional and personal contexts.

Appendix

Puzzle Selection and Categorization

We used a total of 12 puzzles for this study, categorized by difficulty and type. We decided the level of difficulty based on grid size, type of constraints, and whether mathematical operations were involved. The 4 categories are language-based, numerical, spatial, and crossword.

We selected games from The New York Times and BrainBashers because they were the most accessible. These organizations were not involved in or affiliated with this research, and their inclusion here is not an endorsement or promotion of their products or services.

Language-based puzzles require players to find or connect random words based on logical links. There are no levels of gameplay, and the answers can’t be looked up organically.

    • The New York Times’ Wordle [Easy]
        • Goal: Find the 5-letter word of the day.
        • Input: A grid with 5 blank spaces.
    • The New York Times’ Connections [Medium]
        • Goal: Group the terms into the correct mystery categories, forming 4 groups of 4.
        • Input: A 4×4 grid with 16 terms randomly scattered.
    • BrainBashers’ Daily Puzzle [Slightly Harder]
        • Goal: Solve the riddle using reasoning.
Image displaying an example of each of the three language-based puzzles evaluated: Wordle, Connections, and the Daily Puzzle.

Numerical puzzles require the player to arrange a limited set of numbers in a grid so that algebraic constraints are met. The level of gameplay can be adjusted. Most constraints can be met individually with multiple numeric combinations, but only one combination allows for all constraints to be met simultaneously.

    • BrainBashers’ Futoshiki [Easy]
        • Goal: Place numbers 1–4 in the appropriate blank spaces to satisfy the constraints. Each number should appear once per row and once per column.
        • Input: A 4×4 grid with “<” or “>” symbols placed between some of them. 
    • BrainBashers’ MathemaGrids [Medium]
        • Goal: Place numbers 1–9 in the appropriate blank spaces to satisfy the algebraic constraints.
        • Input: A 3×3 grid with 3 random constraints, 6 blank squares, and mathematical symbols in between.
    • BrainBashers’ CalcuDoku [Slightly Harder]
        • Goal: Place numbers 1–4 in the appropriate blank spaces. Each number should appear once per row and once per column. Each cage clue tells you the answer after the cage values have undergone the specified mathematical operation.
        • Input: A blank 4×4 grid with inside “cages” that display a mathematical symbol and a number in the corner.
Image displaying an example of each of the three numerical puzzles evaluated: Futoshiki, MathemaGrids, and CalcuDoku.

Spatial puzzles require the player to color or mark the right squares in a grid so that constraints are met. The level of gameplay can be adjusted.

    • BrainBashers’ 3-In-A-Row [Easy]
        • Goal: Fill the grid with blue (X) or white (O) squares. Every row and column has an equal number of each square, and a 3-in-a-row of the same type isn’t allowed.
        • Input: A 6×6 grid with some spaces already marked.
    • BrainBashers’ Nonogrids [Medium]
        • Goal: Find all of the blocks. The sets of clues are the lengths of the runs, in order, of the blocks in that row or column. Each set must be separated by at least one blank space, and all vertical and horizontal clues must be satisfied simultaneously.
        • Input: A blank 5×5 grid with clues at the beginning of each row and column.
    • BrainBashers’ Daily Tents [Slightly Harder]
        • Goal: Find all of the tents in the forest. Every tent is attached to exactly one tree, and every tree is attached to exactly one tent. The total number of tents is the same as the total number of trees. The clues tell you how many tents are in that row or column. A tent can only be found horizontally or vertically next to a tree, and all vertical and horizontal clues must be satisfied simultaneously.
        • Input: An 8×8 grid with some spaces already marked as “trees” and clues at the beginning of each row and column.
Image displaying an example of each of the three spatial puzzles evaluated: 3-In-A-Row, Nonogrids, and Daily Tents.

Crossword puzzles require the player to find the word that corresponds to the given clue and arrange it in a grid. The level of gameplay can be adjusted, and the answers to each hint can be looked up online.

    • The New York Times’ The Mini [Easy]: 5×5 grid.
    • The New York Times’ The Midi [Medium]: 9×9 grid.
    • The New York Times’ The Crossword [Slightly Harder]: 15×15 grid.
Image displaying an example of each of the three crossword puzzles evaluated: The NYT's Mini, Midi, and Crossword.

Testing Methodology

Our team tested the four AI models using the following settings: 

    • Claude 4.6: Sonnet, Incognito chat.
    • Gemini 3: Fast, temporary chat.
    • ChatGPT 5.3: Fast, temporary chat.
    • Grok 4.1: Fast, private chat.

We asked them to solve the puzzles by uploading a screenshot of each, along with simple instructions on how to answer them.

For Connections, for instance, the prompt used was: “Solve by grouping the terms into the correct mystery categories, forming 4 groups of 4.” For Futoshiki: “Solve by placing numbers 1-4 in the appropriate blank spaces to satisfy the constraints. Each number should appear once per row and once per column.”

For Wordle, we followed the AI model’s suggestions until the right word was found. We noted down each suggestion to see if there’s a pattern to the proposed words.

For the rest of the puzzles, we compared AI models’ outputs to the solution given by the puzzle providers. If the answers were correct, we requested the models to show how they arrived at the solutions. If incorrect, we recorded the mistakes, instructed the models to try again, and noted their responses. Failure on the second try was marked as an overall fail.

Correct answers for puzzles from The New York Times were retrieved from the Try Hard Guides, while answers for BrainBashers puzzles were readily available on their site.

We repeated this process for all the daily puzzles of an entire month (31 days) and logged the following metrics:

    • Average Solve Rate (ASR): The percentage of puzzles for which the model produced the complete and correct final solution grid.
        • ASR (1st): Success rate on the first attempt.
        • ASR (2nd): Success rate on the second attempt (when the first attempt failed).
        • ASR (total): Overall success rate (solved correctly on either the first or second try).
    • Average Solve Time (AST): How long it took for each model to return the correct answers. This only includes successful solves, not time spent on incorrect attempts.
    • Average Number of Tries (ANT): The average number of attempts needed to reach the correct answer (this was only measured for Wordle, where the maximum is capped at 6 guesses).

Limitations of the Study

AI models are nondeterministic systems, meaning their outputs are unpredictable and will rarely (if ever) be the exact same, even when the input parameters are identical. This makes the results of our research both unique and subject to minor trend differences. 

That said, we measured and assessed the most practical and quantifiable metrics to give us room for confident assessments. Within the scope of the puzzles, prompts, and settings we tested, these results offer a useful, if limited, snapshot of how these models handle tasks that require step-by-step reasoning.

Go ahead and reuse the data or images from this article. We don’t mind sharing the details! Just be a good sport and link back to the original article so we get the credit.


References:

  1. When robots can’t riddle: What puzzles reveal about the depths of our own minds—BBC
  2. How GPT‑5 helped mathematician Ernest Ryu solve a 40-year-old open problem—OpenAI
  3. GPT-4 Can’t Reason—arXiv.org
  4. ‘The illusion of thinking’: Apple research finds AI models collapse and give up with hard puzzles—Mashable
  5. Why It Matters That AI Sucks at Sudoku—CNET
  6. How the U.S. Public and AI Experts View Artificial Intelligence—Pew Research Center

Leave a comment

Write a comment

Your email address will not be published. Required fields are marked*