Analysis · 3 August 2026

My view on this project after the third month of data

Grounding is pretty much all that matters for an observation over time, and grounding is what drives the costs up.

Once a month since June I’m capturing the responses to 33 to 35 questions in five categories. I want to understand and record which mobile apps are recommended by the most common AI tools (and Mistral for the love of Europe). Today’s analysis covers grounding (web search), costs, why Perplexity is one of the big players (mentioning-wise) and a teaser for a fitness app deep-dive.

The Headlines:

  • Grounding (web search) matters most for the results
  • Grounding is what drives the costs up
  • Hypothesis: A tool claiming that daily or weekly changes in recommendations by AI tools matter is lying, mentions and recommendations don’t change much
  • AI tools category has the clearest winners with ChatGPT, Claude and Gemini. Perplexity being number four is a sign of great AEO, not a result of user size
  • Fitness apps have great differences in the tools recommending them and several apps are recommended heavily by one AI while not being mentioned at all by another

1. Grounding (web search) matters most for the results

When I first started this project in June I investigated if I really needed grounding, mainly because of the next topic, costs. I ran a few comparisons between grounded and ungrounded and realized grounding is pretty much all that matters for an observation over time. Each model has a cutoff point which is the timeframe until they were trained. After that point in time, the models don’t know what happened in the world anymore. For the current models I use here these dates are:

Model (engine)ReleaseKnowledge cutoffGap at releaseStale by the Aug 2026 run
gpt-5.4-mini (ChatGPT)Mar 17, 2026Aug 31, 2025~6.5 months~11 months
claude-haiku-4-5 (Claude)Oct 15, 2025Feb 2025 “reliable” / Jul 2025 training~8 / ~3.5 months~18 months
gemini-3.1-flash-lite (Gemini)May 7, 2026 (GA)*Jan 2025~16 months*~19 months
mistral-small-2603 (Mistral)Mar 16, 2026Not disclosed by vendorn/aunknown

* Gemini’s gap is measured to the GA date; the model was available in preview before May 2026, so the true gap at release is smaller. The staleness column is unaffected.

Only Google AI, which is the Google AI overview you get when you google (different to Gemini) has no cutoff, it’s live web search.

Anything after these dates above, the LLM doesn’t know and even changes before the cutoff stay wrong in memory, which grounding fixes. Example: ungrounded Mistral recommended “Bard” six times (Bard was renamed to Gemini in 2024). With grounding that number dropped to one mention, correctly phrased as “Gemini, formerly Bard”.

By design, the recommendations are not going to change. Not set in stone because AI responses are probabilistic, but nothing that happens in the world after that date changes their responses fundamentally. And the last two columns demonstrate how severe that gap is - especially for an index that captures changes over time.

2. Grounding is what drives the costs up

completion web_search 8 6 4 2 0 Mistral, June 2026, completion: 0,41 EUR Mistral, June 2026, web_search: 4,49 EUR 4,90 Mistral, July 2026, completion: 0,18 EUR Mistral, July 2026, web_search: 3,60 EUR 3,78 Mistral, August 2026, completion: 0,19 EUR Mistral, August 2026, web_search: 3,60 EUR 3,79 OpenAI, June 2026, completion: 1,15 USD OpenAI, June 2026, web_search: 1,93 USD 3,08 OpenAI, July 2026, completion: 0,85 USD OpenAI, July 2026, web_search: 1,28 USD 2,13 OpenAI, August 2026, completion: 0,80 USD OpenAI, August 2026, web_search: 1,22 USD 2,02 Gemini, June 2026, completion: 0,67 EUR (web search inside free quota) 0,67 Gemini, July 2026, completion: 0,41 EUR (web search inside free quota) 0,41 Gemini, August 2026, completion: 0,22 EUR (web search inside free quota) 0,22 Claude, June 2026, completion: 5,03 USD Claude, June 2026, web_search: 2,05 USD 7,08 Claude, July 2026, completion: 2,75 USD Claude, July 2026, web_search: 1,29 USD 4,04 Claude, August 2026, completion: 1,90 USD Claude, August 2026, web_search: 1,36 USD 3,26 Jun Jul Aug Jun Jul Aug Jun Jul Aug Jun Jul Aug Mistral (EUR) OpenAI (USD) Gemini (EUR) Claude (USD)
Cost per monthly run per engine, stacked into completion and web_search. Mistral and Gemini bill in EUR, OpenAI and Claude in USD, no conversion applied, so compare months within an engine, not engines against each other. Gemini shows no web_search share because my volume sits inside the free search quota.
EngineJune 2026July 2026August 2026
Mistral (EUR)completion0,410,180,19
web_search4,493,603,60
OpenAI (USD)completion1,150,850,80
web_search1,931,281,22
Gemini (EUR)completion0,670,410,22
web_search000
Claude (USD)completion5,032,751,90
web_search2,051,291,36

(Gemini web search is inside my free quota (5,000 free search requests per month), so my costs are tokens only.)

This chart visualizes that overall the main cost driver is web search, with Claude as an exception. I chose to use the small-tier models because I didn’t see the upside over frontier models (i.e. going from Haiku to Sonnet or Opus on Claude, or using ChatGPT not-mini); the questions are simple and require no complex problem solving skills or deep thinking. And I chose to run each question in the set of 33 to 35 for one category only once. I tested that with three runs in June with one category (fitness) and found that the differences in multiple runs are not drastic enough to justify a cost multiplier by 3 or more. The results here are directional. More details on that matter here: /notes/one-sample-is-enough/.

The message here is that if you are paying for a tool that monitors your brand ranking in AI tools, pay attention if they are measuring the mentions with grounding enabled. This also gives you a very rough overview of the costs vs. the price. A monthly run with 5 categories x 4 AI tools (excl. Google AI) x 33 to 35 questions costs me around 8 Euro. That gives me 688 grounded answers a month (172 questions across 4 tools). If that’s the scale you look at in an AEO monitoring tool daily, monthly costs for them are around 30 x 8 Euro = 240 Euro.

Which now leads me to my next topic.

3. A tool claiming that daily or weekly changes in recommendations by AI tools matter is lying

In the third run I’m still waiting for a change at the top spot in all categories. Not in a single category the top app changed. I see some changes in the scores and in the ranks below the top spot, and possibly there are trends, like Adobe Express declining (21 -> 18 -> 11) and Caliber climbing (15 -> 19 -> 23) that will lead to position changes, so I will keep an eye on this to call it out as soon as this happens, but my stance stands that monitoring this daily makes no sense.

Running a big round might make sense with model updates, but even that has to be seen (see my first point). For me, for now, monthly is a good check-in cadence.

4. Perplexity being number four in the AI apps category is a sign of great AEO, not a result of user size

AppIndex score (Aug)Best usage figureAs ofWho says
ChatGPT46 (#1)900M weekly active usersFeb 2026OpenAI self-reported
Claude41 (#2)~245M MAU (web + mobile app)Q2 2026Sensor Tower estimate; Anthropic publishes nothing
Gemini34 (#3)950M+ MAU (Gemini app)Jul 2026 earningsGoogle self-reported
Perplexity32 (#4)100M+ MAU “across all products”~Apr 2026CEO self-reported (via Sacra)
Microsoft Copilot19 (#5)No standalone number exists; 1.6% of global AI-assistant app shareMay 2026Sensor Tower
Grok1117M MAU “Grok AI features” inside XMar 2026SpaceX S-1 filing
Meta AI11B+ MAU bundled across FB/IG/WhatsAppMay 2025, never updatedZuckerberg
DeepSeek3127–130M MAU (China-centric panel)Mar–May 2026QuestMobile estimate
Qwen app4100M self-reported / 166M QuestMobile; app is China-onlyJan / May 2026Alibaba / QuestMobile
Le Chat5No verified figure newer than Feb 2025 (1M downloads in 14 days, “several million” regular users)Feb 2025Mistral CEO via Reuters

Gemini’s monthly active users are inflated by Android users. Meta AI claims a billion users and without knowing how many of the Facebook, Instagram, WhatsApp & co users actually use the Meta AI, it probably deserves a higher score than 1 or rank number 12 in the index. Same with Grok on X. But my set of questions is more after a standalone AI assistant, that’s to their disadvantage.

Qwen and DeepSeek have real 100 million+ scale on China-focused panels and my en-US index barely registers them.

My take is that Perplexity is the most over-recommended app in the dataset. They are relevant, they are persistently up there over 3 months and they are not small, but nowhere near as big as ChatGPT or Gemini, nowhere near as big as Claude. So how does Perplexity end up in the top 4? As far as the data can tell, Perplexity is hyped within the AI bubble and the tech press writes about it a lot. And grounded models read the tech press, so they recommend it. I’d conclude that they are good at AEO. And that makes them deserve their spot in the top 4.

Fitness is the most volatile category measured by median month-to-month change, it’s 4 points versus 1 in AI. Fitness has 13 apps with a score of 5 or higher, AI has only 6.

The app that ranks number 1 in the Health and Fitness app store category on Apple (US) is Strava. It ranks at number 5 on Google Play. Sources: https://apps.apple.com/us/iphone/charts/6013 and https://play.google.com/store/apps/category/HEALTH_AND_FITNESS?gl=US, data from 2026-08-03.

Nike Training Club, top recommended app by AI tools, is on rank 104 in that category (US, iPhone). Hevy ranks 25, Fitbod 19, Caliber 132, Strong 149.

I will deep dive here next and publish a standalone analysis.

The numbers here come from the index. Every category has its own leaderboard and full table, and the method is published in full.