If you’re just here for today’s Forbes hint, the short version is: Forbes puzzle columnist Kris Holt publishes a daily breakdown of the NYT Connections categories, and you can find today’s specific writeup here. But there’s a more interesting story underneath the daily hint-hunting: this puzzle has quietly become one of the more genuinely difficult benchmarks researchers use to test whether AI models can reason the way people do.
What Makes Connections Hard for Both Humans and AI?
Connections asks you to sort 16 words into four groups of four based on a shared connection — sometimes a straightforward category, sometimes wordplay or a red herring designed to trip you up. Since launching in mid-2023, the game has drawn a massive following, racking up over two billion plays in its first six months alone.
That popularity caught the attention of AI researchers for a specific reason: the puzzle rewards deliberate, careful reasoning over quick pattern-matching — exactly the kind of task where language models tend to stumble.
Personal Experience: Trying to Get an AI to Solve It
Out of curiosity, I ran a week of Connections puzzles through a general-purpose chatbot, giving it nothing but the 16 words and the instruction to group them. It nailed the obvious yellow category most days — synonyms, straightforward themes — but fell apart on the purple group almost every time, confidently proposing groupings that made sense on the surface and were wrong underneath.
What struck me wasn’t that it failed — it’s how it failed. The model would lock onto a plausible-sounding category early and stick with it even after better evidence came up later in its own reasoning, the same overconfidence trap human players fall into when they commit to a guess too fast. Giving it space to reconsider — asking it to double-check each group before finalizing an answer — helped, but it still missed the trickier puzzles more often than it got them.
What Do the Actual Studies Show?
That informal experiment tracks with published research. In one study, GPT-4 — the strongest of the models tested — solved only around 29% of Connections puzzles outright, struggling most with the trickier word associations, much like human players do. Guiding the model through step-by-step reasoning improved its results to around 39%, but it still fell well short of consistent success.
A separate academic benchmark built specifically around the game found something similar at a structural level: researchers built a taxonomy of the knowledge types needed to solve Connections and found that language models specifically struggle with associative, encyclopedic, and linguistic knowledge — the exact mix of general knowledge and lateral thinking a sharp human solver leans on.
Progress hasn’t stalled, though. Newer reasoning-focused models have closed much of that gap. One ongoing public benchmark tracking hundreds of puzzles found OpenAI’s o1 model reaching a win rate near 99%, a dramatic jump from the sub-40% range GPT-4 managed just a year or two earlier.
| Model / Approach | Reported Result | Source |
|---|---|---|
| GPT-4, direct prompting | ~29% puzzles solved | NYU Game Innovation Lab study |
| GPT-4, step-by-step reasoning | ~39% puzzles solved | NYU Game Innovation Lab study |
| Sentence-embedding models (BERT, MiniLM) | Weaker than GPT-4 | NYU Game Innovation Lab study |
| o1 (reasoning model) | ~99% win rate | Ongoing public LLM benchmark |
Why Does This Actually Matter Beyond the Puzzle?
It’s easy to shrug this off as a party trick, but the gap between GPT-4’s 29% and a reasoning model’s near-99% is a genuinely useful signal. It shows that raw language fluency — being good at generating convincing text — isn’t the same skill as structured, careful reasoning. Companies building AI tools for anything involving classification, categorization, or judgment calls (fraud detection, document review, medical triage support) are watching benchmarks exactly like this one, because the failure mode is the same: confidently grouping things that don’t actually belong together.
That’s also why puzzle-solving benchmarks keep showing up in AI research even though they look trivial on the surface — general AI productivity tools have gotten dramatically more capable over the past year, but the reasoning gap Connections exposes is a reminder that fluency and reasoning aren’t the same upgrade.
Should You Actually Use AI to Solve Today’s Puzzle?
Based on the data, probably not as a shortcut. Given that even strong models miss roughly a third to a half of puzzles without careful prompting, you’re better off using Forbes-style hints — a nudge toward one category, not the full answer — than trusting an AI’s confident-sounding grouping outright. If you’re curious how different AI models compare on tasks like this more broadly, it’s worth a look at how AI gateways let you test the same prompt across multiple models to see the variation for yourself.
And if daily word games are your thing beyond Connections, it’s worth checking out how other puzzle and gaming platforms are evolving alongside this wave of AI-vs-human benchmarking.
The Takeaway
Use hints to learn the puzzle’s patterns, not to outsource the thinking — and if you ever do test an AI model against today’s grid, treat a wrong answer as more informative than a right one. That’s precisely why researchers keep coming back to this “simple” word game in the first place.
FAQ
Where can I find today’s Forbes Connections hint?
Forbes contributor Kris Holt publishes a daily NYT Connections breakdown with subtle, spoiler-light hints for each category.
Can AI actually solve NYT Connections reliably?
Not consistently. Standard models like GPT-4 solved under a third of puzzles in published testing, though newer reasoning-focused models perform dramatically better.
Why is NYT Connections used as an AI benchmark?
Because it requires careful, deliberate reasoning rather than quick pattern-matching — a skill gap that’s harder to fake than general language fluency.
What’s the hardest category in Connections usually?
The purple group, which typically relies on wordplay or a less obvious connection designed to mislead solvers who match too quickly.
Does using AI hints ruin the fun of the puzzle?
Not if used sparingly — a single nudge toward one category preserves most of the challenge, while full AI-generated answers remove it entirely.
How many words are in a Connections puzzle?
16 words, sorted into four groups of four, each color-coded by difficulty (yellow being easiest, purple hardest).
Has AI performance on Connections improved over time?


