Methodology

How Much Do AI Answers Change From Day to Day? Measuring the Noise Before You Trust a ‘Before / After’

Ask ChatGPT, Perplexity, Gemini or Claude the same buying question three times today, then again tomorrow. The names in the answer shift, and the sources they cite shift more. Before anyone shows you a “before / after” for a fix, you should know how much the engines move on their own. This page explains how CitePulse measures that noise, publishes it, and reads every result against it.

The short answer

A lot. Published measurements from 2026 put the day-to-day turnover of the sources AI engines cite at roughly two thirds: run the same question on consecutive days and about a third of the cited URLs are the same. Repeat one question three times on the same day and the three answers agree only partially on whether a given brand is cited at all. A single run is therefore a snapshot, not a state.

Nobody in the AI-visibility tool market publishes this number for their own instrument. CitePulse does, daily, at citepulse.ai/methodology and as machine-readable JSON at api.citepulse.ai/api/public/noise.

Where the numbers come from

  • Same model name, different instrument. A preregistered study of 52,988 queries found that identical prompts sent to the same model on consecutive days agreed on rankings at a Spearman correlation of about 0.78, not the 0.99 you would expect from a fixed instrument (arXiv:2609.04198). On a shared endpoint the model name is not a frozen measuring device.
  • Repetitions disagree. Re-asking the same question yields cited-source sets with a Jaccard overlap of only 0.50–0.61 (arXiv:2605.27440).
  • Sources turn over daily. About 65% of cited sources change from one day to the next (arXiv:2604.07585).
  • Citations decay on their own. Once a page is cited, that citation fades with a half-life of roughly 39–68 days even if nothing changes (arXiv:2607.15771). An old “before” is not a baseline.

What CitePulse measures every day: the instrument panel (NOISE-1.0)

We keep a fixed panel: five frozen buying questions about our own site, citepulse.ai, asked on four engines (Perplexity, Gemini, ChatGPT, Claude), three times each, once a day. Sixty calls, the same bytes, every day. From those rows we compute, per engine:

  • Noise band: the median absolute change in citation share per question, day over day. This is the number a “before / after” is compared against.
  • Repetition agreement: the share of question-days on which all three repetitions gave the same cited / not-cited result.
  • Source overlap day to day: the median Jaccard similarity of the cited-URL set for a question between consecutive days.
  • Days measured: how many days of data the band rests on. Below 14 days an engine gets a default band of 0.34 (more than one repetition in three), and the interface says so.

None of this enters any score. It exists so that a client reading their own results knows how jumpy the ruler is.

How a fix is read against the band

When a client marks a change as done in their CitePulse dashboard (a schema file, a new answer page, a pricing page rewrite), the change gets a date and a baseline: up to three earlier runs of the same frozen question set and the same method version, none older than 60 days. Later monitoring runs are summed per question and engine on the “after” side. Each cell then gets one of three states:

  • Rose beyond the band: the after-rate minus the before-rate is larger than that engine’s noise band.
  • Fell beyond the band: the drop is larger than the band.
  • Within normal variation: everything else, which is where most cells land, and which is an honest result, not a failure.

A result is reported as observed only with at least two comparable runs, at least six repetitions per cell on the after side, at least 14 days since the change, and a baseline no older than 60 days. Fewer than that and the dashboard says “measuring” and how long until the next run.

What this is not

It is not proof that a change caused anything. Live engines do not allow that claim, and we do not make it. The wording throughout CitePulse is “after this change” and “since 2026-09”, never “thanks to this change”. What you get is an observation with a date, an engine and a noise band, which is more than a screenshot of one lucky answer.

“Seen in k of n runs”

The same logic applies to your own pages. Instead of “AI cites this page”, the dashboard says in how many of the last comparable runs a page appeared among the cited sources, for example 2 of 3. Runs are comparable only when they used the same frozen questions and the same measurement version; a change of method starts a new series rather than mixing rulers.

Frequently asked questions

How much do AI answers change from day to day?

Roughly two thirds of cited sources turn over day to day on the same question, and same-day repetitions agree only partially. The live numbers for our own panel are on the methodology page and refresh daily.

What is a noise band and why does a before/after need one?

The noise band is the median day-over-day change in citation share per question on a fixed panel, per engine. A change smaller than the band is within normal variation and is reported as such. Until an engine has 14 days of panel data the band is a stated default of 0.34.

What does “seen in k of n runs” mean?

In how many of the last comparable runs a page appeared among the cited sources. A single run is a snapshot; k of n is a measurement.

Can you prove a fix caused more AI citations?

No. CitePulse reports whether citations moved beyond the noise band after a dated change, per question and engine, with at least two runs and 14 days. That is correlation with a date and a band, never proof of cause.

Related reading: how to check if AI cites your brand, citation share as the new share of voice, and the full CitePulse methodology, including the live NOISE-1.0 numbers. Run a free audit to see your own pages “seen in k of n”.