Your analytics stack can already tell you that AI traffic exists. GA4 shows sessions with chatgpt.com in the referrer, PostHog can chart them by week, and market-level counters keep publishing growth curves — Adobe measured a 1,324% rise in AI-referred traffic to retail sites since 2024. What none of those tools can tell you is which product page the assistant recommended, what the visitor did next, and whether that line in your dashboard was a human, a bot, or a scraper wearing a browser. That gap has a name: it is the difference between traffic analytics and attribution.
We run both kinds of measurement on this very blog, so this is not a hit piece on analytics tools. It is a map of where each one stops.
The camera and the receiving dock
A security camera at the front door of a store sees that someone walked in, roughly when, and which direction they came from. That is what traffic analytics does, and it does it well. The receiving dock at the back of the same store works differently: every item that arrives gets scanned, logged against a purchase order, and signed for. Nobody reconciles inventory from camera footage.
AI traffic is now big enough to deserve the receiving-dock treatment. Salesforce projects $262 billion in AI-driven commerce for 2026 — around 20% of online sales. When a fifth of the channel flows through assistants, "we saw some visits from chatgpt.com" stops being a sufficient level of detail, for the same reason "some boxes arrived on Tuesday" was never an inventory system.
What your current analytics already do well
Honesty first: a referrer report in GA4 or PostHog is genuinely useful, and it is where every store should start. It will show you whether AI-referred sessions exist at all, which assistant sends more of them, and how the trend moves month over month. PostHog will additionally give you session-level behavior — pages per visit, clicks, drop-off — for the AI-referred segment, the same way it does for any other segment. If you have not yet segmented your traffic by AI referrers, do that this week; our tracking guide for Shopify stores walks through the exact setup and the signals worth watching.
The catch is that everything these tools know about an AI visit is inherited from a generic web-analytics model that predates AI assistants. The referrer string is the entire identity of the visit. What the assistant actually said, which product it pointed the shopper at, and whether the "visitor" was a human at all — none of that is in the referrer.
Three questions a counter cannot answer
Which product did the assistant recommend? An AI assistant does not send traffic to "your site." It sends a shopper to a specific page because of a specific claim it believed about a specific product. Store-level referral counts flatten this into one number, which makes the most important question — which products win AI recommendations and which are invisible — unanswerable. The difference matters commercially: Adobe measured AI-referred retail visitors converting 54% more than average in May 2026, with 37% higher revenue per visit. Numbers like those, applied blindly at store level, tell you to celebrate. Applied at product level, they tell you what to fix.
Was it a human? AI-era traffic comes in at least three distinct kinds: human shoppers referred by an assistant, crawlers reading your pages to build the answers, and automated agents acting on someone's behalf. A generic analytics pageview lumps whichever of them execute JavaScript, and silently misses most crawlers, which do not. These populations mean opposite things. Crawler visits are supply — evidence that assistants are ingesting your catalog, worth tracking on its own (our GEO playbook explains why crawl access is layer one). Human referrals are demand. Mixing them in one metric corrupts both.
Can you defend the data? Third-party analytics collect first and let you configure compliance later. Attribution-grade measurement inverts this: a privacy decision — consent state, regional posture, Global Privacy Control — is evaluated before a row is written, and a visit that fails the check is never recorded rather than recorded and filtered. If AI-attributed revenue is going to appear in board decks and merchandising decisions, the lineage of every row has to survive scrutiny. "We think the pixel was configured correctly that quarter" does not.
What attribution-grade measurement looks like
The artifact is different from a dashboard. It is a first-party ledger — collected on your own domain, owned by you — where each row is classified at the moment of collection: which AI surface referred the visit (a maintained registry of assistant referrers and their URL signatures, not a regex you wrote once), which page and product it landed on, and a session linkage that lets "landed from Perplexity" connect to "added to cart" without third-party cookies. Crawler traffic is captured separately at the server, where JavaScript-free agents are actually visible, so the human ledger stays human.
This is also why conversion benchmarks from the market — Similarweb's 7.1% for AI referrals against 7.8% for organic search, Adobe's +54% — should be treated as invitations to measure, not as answers. Both numbers are averages over stores whose AI visibility ranges from excellent to nonexistent. Your store's number depends on what the assistants can actually see of your catalog, which is exactly the thing you control. Across 800+ stores we have analyzed, the dominant pattern is a store that looks fine on brand queries and returns zero on buyer queries — a profile no market-average conversion rate describes. If you have never measured your own baseline, start with an AI visibility audit; it is the diagnostic this whole loop hangs off.
Measurement is a third of the job
Here is the structural difference between Arbling and the analytics and monitoring category (we compared the monitoring tools honestly in this guide): most tools in this space measure, report, and stop. The full path has three legs. Measure — audit what assistants say about you, with repeated sampling, on a frozen panel so month-over-month deltas are real. Repair and republish — fix the structured product data the assistants failed on and push it back to the surfaces agents actually read. Attribute — run the receiving-dock ledger so the revenue effect of the repair shows up as rows, not as a feeling. To our knowledge we are the only vendor operating all three legs as one loop; every alternative we respect does one of them well.
A measurement number that never feeds a repair is trivia. A repair that never feeds back into attribution is an invoice you cannot justify. The loop is the product.
We run all three layers on this blog
The page you are reading is instrumented three ways, and the layers deliberately do not overlap. GA4 and PostHog run as the camera: worldwide, aggregate, session-level human analytics. A server-side crawler monitor logs the JavaScript-free agents — the assistants' own bots reading these guides. And our own attribution beacon runs as the receiving dock: first-party, privacy-gated (it honors Global Privacy Control and ships with a visible notice and a working opt-out in the footer below), classifying human AI-referrals against the same referrer registry our merchant product uses.
We do this because we sell the beacon, and eating our own instrumentation means an integration mistake bites us before it bites a customer. It already has: our own dogfood install surfaced an integration footgun — a beacon served from the wrong origin posts its data into a void — that is now caught before any merchant can reproduce it. That is the kind of bug a vendor only finds by running their own product in production.
When a counter is enough
If you are pre-launch, if AI referrals are still a rounding error in your channel mix, or if you only need to know whether the trend exists — a referrer segment in the analytics you already own is the right tool, and adding attribution machinery would be premature. The moment to graduate is when you start making merchandising, content, or ad-budget decisions based on what AI assistants do with your catalog. Decisions need product-level, human-verified, defensible rows. Cameras were never built for that.
The first step is free either way: run the AI Readiness Score on your store and see what the assistants can currently see. The measurement takes minutes. What you do with it is the part that moves revenue.