Analysis · 3 August 2026
My view on this project after the third month of data
Grounding is pretty much all that matters for an observation over time, and grounding is what drives the costs up.
Once a month since June I’m capturing the responses to 33 to 35 questions in five categories. I want to understand and record which mobile apps are recommended by the most common AI tools (and Mistral for the love of Europe). Today’s analysis covers grounding (web search), costs, why Perplexity is one of the big players (mentioning-wise) and a teaser for a fitness app deep-dive.
The Headlines:
- Grounding (web search) matters most for the results
- Grounding is what drives the costs up
- Hypothesis: A tool claiming that daily or weekly changes in recommendations by AI tools matter is lying, mentions and recommendations don’t change much
- AI tools category has the clearest winners with ChatGPT, Claude and Gemini. Perplexity being number four is a sign of great AEO, not a result of user size
- Fitness apps have great differences in the tools recommending them and several apps are recommended heavily by one AI while not being mentioned at all by another
1. Grounding (web search) matters most for the results
When I first started this project in June I investigated if I really needed grounding, mainly because of the next topic, costs. I ran a few comparisons between grounded and ungrounded and realized grounding is pretty much all that matters for an observation over time. Each model has a cutoff point which is the timeframe until they were trained. After that point in time, the models don’t know what happened in the world anymore. For the current models I use here these dates are:
| Model (engine) | Release | Knowledge cutoff | Gap at release | Stale by the Aug 2026 run |
|---|---|---|---|---|
| gpt-5.4-mini (ChatGPT) | Mar 17, 2026 | Aug 31, 2025 | ~6.5 months | ~11 months |
| claude-haiku-4-5 (Claude) | Oct 15, 2025 | Feb 2025 “reliable” / Jul 2025 training | ~8 / ~3.5 months | ~18 months |
| gemini-3.1-flash-lite (Gemini) | May 7, 2026 (GA)* | Jan 2025 | ~16 months* | ~19 months |
| mistral-small-2603 (Mistral) | Mar 16, 2026 | Not disclosed by vendor | n/a | unknown |
* Gemini’s gap is measured to the GA date; the model was available in preview before May 2026, so the true gap at release is smaller. The staleness column is unaffected.
Only Google AI, which is the Google AI overview you get when you google (different to Gemini) has no cutoff, it’s live web search.
Anything after these dates above, the LLM doesn’t know and even changes before the cutoff stay wrong in memory, which grounding fixes. Example: ungrounded Mistral recommended “Bard” six times (Bard was renamed to Gemini in 2024). With grounding that number dropped to one mention, correctly phrased as “Gemini, formerly Bard”.
By design, the recommendations are not going to change. Not set in stone because AI responses are probabilistic, but nothing that happens in the world after that date changes their responses fundamentally. And the last two columns demonstrate how severe that gap is - especially for an index that captures changes over time.
2. Grounding is what drives the costs up
| Engine | June 2026 | July 2026 | August 2026 | |
|---|---|---|---|---|
| Mistral (EUR) | completion | 0,41 | 0,18 | 0,19 |
| web_search | 4,49 | 3,60 | 3,60 | |
| OpenAI (USD) | completion | 1,15 | 0,85 | 0,80 |
| web_search | 1,93 | 1,28 | 1,22 | |
| Gemini (EUR) | completion | 0,67 | 0,41 | 0,22 |
| web_search | 0 | 0 | 0 | |
| Claude (USD) | completion | 5,03 | 2,75 | 1,90 |
| web_search | 2,05 | 1,29 | 1,36 |
(Gemini web search is inside my free quota (5,000 free search requests per month), so my costs are tokens only.)
This chart visualizes that overall the main cost driver is web search, with Claude as an exception. I chose to use the small-tier models because I didn’t see the upside over frontier models (i.e. going from Haiku to Sonnet or Opus on Claude, or using ChatGPT not-mini); the questions are simple and require no complex problem solving skills or deep thinking. And I chose to run each question in the set of 33 to 35 for one category only once. I tested that with three runs in June with one category (fitness) and found that the differences in multiple runs are not drastic enough to justify a cost multiplier by 3 or more. The results here are directional. More details on that matter here: /notes/one-sample-is-enough/.
The message here is that if you are paying for a tool that monitors your brand ranking in AI tools, pay attention if they are measuring the mentions with grounding enabled. This also gives you a very rough overview of the costs vs. the price. A monthly run with 5 categories x 4 AI tools (excl. Google AI) x 33 to 35 questions costs me around 8 Euro. That gives me 688 grounded answers a month (172 questions across 4 tools). If that’s the scale you look at in an AEO monitoring tool daily, monthly costs for them are around 30 x 8 Euro = 240 Euro.
Which now leads me to my next topic.
3. A tool claiming that daily or weekly changes in recommendations by AI tools matter is lying
In the third run I’m still waiting for a change at the top spot in all categories. Not in a single category the top app changed. I see some changes in the scores and in the ranks below the top spot, and possibly there are trends, like Adobe Express declining (21 -> 18 -> 11) and Caliber climbing (15 -> 19 -> 23) that will lead to position changes, so I will keep an eye on this to call it out as soon as this happens, but my stance stands that monitoring this daily makes no sense.
Running a big round might make sense with model updates, but even that has to be seen (see my first point). For me, for now, monthly is a good check-in cadence.
4. Perplexity being number four in the AI apps category is a sign of great AEO, not a result of user size
| App | Index score (Aug) | Best usage figure | As of | Who says |
|---|---|---|---|---|
| ChatGPT | 46 (#1) | 900M weekly active users | Feb 2026 | OpenAI self-reported |
| Claude | 41 (#2) | ~245M MAU (web + mobile app) | Q2 2026 | Sensor Tower estimate; Anthropic publishes nothing |
| Gemini | 34 (#3) | 950M+ MAU (Gemini app) | Jul 2026 earnings | Google self-reported |
| Perplexity | 32 (#4) | 100M+ MAU “across all products” | ~Apr 2026 | CEO self-reported (via Sacra) |
| Microsoft Copilot | 19 (#5) | No standalone number exists; 1.6% of global AI-assistant app share | May 2026 | Sensor Tower |
| Grok | 1 | 117M MAU “Grok AI features” inside X | Mar 2026 | SpaceX S-1 filing |
| Meta AI | 1 | 1B+ MAU bundled across FB/IG/WhatsApp | May 2025, never updated | Zuckerberg |
| DeepSeek | 3 | 127–130M MAU (China-centric panel) | Mar–May 2026 | QuestMobile estimate |
| Qwen app | 4 | 100M self-reported / 166M QuestMobile; app is China-only | Jan / May 2026 | Alibaba / QuestMobile |
| Le Chat | 5 | No verified figure newer than Feb 2025 (1M downloads in 14 days, “several million” regular users) | Feb 2025 | Mistral CEO via Reuters |
Gemini’s monthly active users are inflated by Android users. Meta AI claims a billion users and without knowing how many of the Facebook, Instagram, WhatsApp & co users actually use the Meta AI, it probably deserves a higher score than 1 or rank number 12 in the index. Same with Grok on X. But my set of questions is more after a standalone AI assistant, that’s to their disadvantage.
Qwen and DeepSeek have real 100 million+ scale on China-focused panels and my en-US index barely registers them.
My take is that Perplexity is the most over-recommended app in the dataset. They are relevant, they are persistently up there over 3 months and they are not small, but nowhere near as big as ChatGPT or Gemini, nowhere near as big as Claude. So how does Perplexity end up in the top 4? As far as the data can tell, Perplexity is hyped within the AI bubble and the tech press writes about it a lot. And grounded models read the tech press, so they recommend it. I’d conclude that they are good at AEO. And that makes them deserve their spot in the top 4.
5. Fitness apps have great differences in the tools recommending them and several apps are recommended heavily by one AI while not being mentioned at all by another
Fitness is the most volatile category measured by median month-to-month change, it’s 4 points versus 1 in AI. Fitness has 13 apps with a score of 5 or higher, AI has only 6.
The app that ranks number 1 in the Health and Fitness app store category on Apple (US) is Strava. It ranks at number 5 on Google Play. Sources: https://apps.apple.com/us/iphone/charts/6013 and https://play.google.com/store/apps/category/HEALTH_AND_FITNESS?gl=US, data from 2026-08-03.
Nike Training Club, top recommended app by AI tools, is on rank 104 in that category (US, iPhone). Hevy ranks 25, Fitbod 19, Caliber 132, Strong 149.
I will deep dive here next and publish a standalone analysis.
The numbers here come from the index. Every category has its own leaderboard and full table, and the method is published in full.