Decide what counts before the work starts
The measurement conversation almost always happens in the wrong order. Work begins, months pass, and then somebody asks whether it is working — at which point everyone goes looking for a number that answers the question favorably, and finds one, because there are always several available.
Agreeing the measures in advance costs an hour and removes that entire failure mode. Three things need to be written down before anything ships: the primary measure that the engagement is judged on, the leading indicators that will move first and are therefore how you tell progress from noise, and the known limits of both — what these instruments cannot see, stated while nobody has an interest in the answer.
That last item is the one that gets skipped, and it is the one that protects everybody. Attribution understates organic search. Search Console cannot show you every query. Rank tracking measures a synthetic query rather than a person. Writing those down in month one converts them from excuses in month nine into shared understanding from the start.
The baseline, and the export that expires
A baseline is not a screenshot of a dashboard. It is a dated file, stored somewhere that survives staff changes, containing the numbers you will later be asked to compare against.
Capture, at minimum: clicks, impressions, average position and click-through rate segmented by brand and non-brand and by template; conversions and revenue by landing page and template; the indexed URL count against the crawled 200 count; Core Web Vitals field data at the 75th percentile split by device; and a note of what was already known to be broken.
One constraint makes this urgent rather than merely advisable. Search Console retains sixteen months of performance data, and both the interface and its spreadsheet export cap at 1,000 rows. Data outside that window is gone, and on any site of size the 1,000-row cap means a single export is a sample rather than a record. Export regularly, by segment, and keep the files. The comparison you will most want in eighteen months is the one that becomes impossible if nobody exported anything today.
Date the file and record what was happening at the time — a promotion, an outage, a seasonal peak, a release. A baseline without context gets misread by whoever inherits it.
Four instruments, four blind spots
Every measurement dispute I have seen comes from asking one instrument a question it cannot answer. Keep the boundaries hard.
- Search Console reports clicks, impressions, click-through rate and average position. It is the only source that sees impressions or position at all. It cannot see what happened after the click, and it cannot show you every query.
- Analytics reports sessions, engagement, conversions and revenue. It never sees an impression or a position. Consent controls and tracking prevention mean measured sessions understate real sessions, and last-click models understate organic's contribution to conversions credited elsewhere.
- Rank tracking reports a position for a synthetic query from a chosen location and device. Results are personalized, localized and device-dependent, so it is a sample, not a universal position. It is a diagnostic instrument, not an outcome metric.
- Share of voice and AI visibility tools report estimates built on the vendor's own keyword universe or prompt set. Two vendors will disagree because they are measuring different sets. For AI surfaces specifically, no vendor has click data, because Google does not provide it — so any vendor reporting AI-driven traffic is modeling, not measuring.
A report that presents all four as one blended score is reporting a tool's opinion. Keep them separate and name what each one is for.
Segment before you total
Site-wide totals hide almost everything worth knowing, and they are the reason so many reports feel simultaneously true and useless.
Four splits do most of the work. Brand versus non-brand is the first and the most important: branded search largely measures your marketing elsewhere, and leaving it in a search report lets a television campaign or a funding announcement look like an SEO result. By template — product, category, article, homepage — because fixes are template-level and so is failure. By device, because mobile and desktop behave differently and are assessed differently for performance. By country or language where more than one exists, because a problem confined to one market is a completely different investigation.
Segmentation also changes what a movement means. A 10% site-wide decline that is entirely one template is a specific, fixable event. The same decline spread evenly across everything is a different story with different causes. You cannot tell those apart from a total, and the total is what most reports lead with.
Cadence: what to look at, and how often
Looking at everything constantly produces reaction to noise. A fixed cadence produces decisions.
Weekly, ten minutes. Anything binary and urgent: Manual Actions and Security Issues reports, Crawl Stats host status, the indexed count's direction of travel, and server error rates. These are checks for faults, not for progress, and none of them should generate a discussion in a normal week.
Monthly. Clicks, impressions and position by segment against the baseline; conversions and revenue from organic landing pages; the state of the work — what shipped, what is queued, what is blocked. A monthly report should be readable in five minutes and should say what changed and what it means, including when the answer is that nothing meaningful changed.
Quarterly. The direction question. Is the primary measure moving, are the leading indicators consistent with it, and is the plan still the right plan. This is the meeting where continuing or stopping is actually decided, which is why it needs the baseline and not a mood.
Resist reporting on cycles shorter than the thing being measured. Field performance data accumulates from real visits. Indexing responds to recrawls. Ranking changes are not instant. A weekly report on any of those is a weekly invitation to overreact.
Read movement against dated events, not against last month
Keep a single annotated timeline: every release, content change, redirect batch, tracking change, campaign start and stop — and every confirmed Google update with its start date and duration, from Google's published ranking release history.
The dates matter more than most people expect. In 2026 the March core update began on 24 March and completed in under twenty hours, and a separate spam update began three days later and ran twelve days. Anyone reading a late-March change without dated data will merge two distinct events into one story and act on the wrong one. The May 2026 core update ran from 21 May for just under twelve days, and the June 2026 spam update ran two days from 24 June.
The test for attributing anything to an update is unglamorous: did the movement begin when the rollout began, and stabilize when it completed. If it started three weeks earlier, the update is not the explanation, however convenient.
Seasonality gets the same treatment. Compare like periods — the same weeks a year earlier, the same number of days, the same weekday composition — and compare against your non-search channels. If everything softened together, the cause is demand, and no amount of search work is going to be credited or blamed correctly until that is said out loud.
What belongs in the report
A report exists so that a decision can be made. Anything in it that cannot change a decision is decoration.
- What changed, in the primary measure, against the dated baseline and the same period last year.
- Why, with the evidence and the alternative explanations that were ruled out. This is the section that separates analysis from a dashboard export.
- What shipped since the last report, and what is blocked and by whom. Implementation queues are usually the binding constraint and should be visible to the person who can unblock them.
- The leading indicators — coverage, non-brand impressions by template, field performance — with a note on whether they are consistent with the primary measure or contradicting it.
- What this data cannot tell you. One short paragraph, every time. Not a disclaimer: a specific statement, such as which segments fell below the row cap or where tracking is known to be incomplete.
- The decision being asked for. Continue, change, stop, or escalate.
What does not belong: a single composite score, keyword position screenshots chosen after the fact, activity counts presented as outcomes, and any number without the window and the segment it came from.
Deciding whether it is working, and when to stop
Set the review point in advance and hold to it. A quarter is usually the shortest window in which a direction is legible for organic work, and the honest framing is that different work runs on genuinely different clocks: a technical fix can take effect within days of a recrawl, a migration takes weeks for a small to medium site by Google's own estimate, and content and authority work is measured in months.
Judge on the leading indicators first, because they move before the outcome does. If coverage improved, non-brand impressions by template rose, and the pages that were fixed are being crawled and stored, the mechanism is working even if revenue has not yet followed. If none of those moved after the work was recrawled and given time, that is a signal to change approach rather than to wait longer.
Be equally clear about what nobody can promise. Google states directly that there is no guarantee that changes you make to your website will result in noticeable impact in search results, and that if there is more deserving content it will continue to rank well. Anyone replacing that with a date and a percentage is not better informed than Google; they are making a sale. Honest measurement is not pessimism — it is the only thing that makes the good numbers believable when they arrive.
Frequently Asked Questions
What should I actually measure to know if SEO is working?
One primary measure and two or three leading indicators, agreed before the work starts. The primary measure should be commercial — revenue or qualified leads from organic landing pages, compared against a dated baseline. The leading indicators are the things that move first: coverage, meaning the count of intended URLs actually indexed; non-brand impressions by template; and field performance where speed work is in scope. Leading indicators let you tell progress from noise months before the commercial measure is legible, which is what keeps a reasonable program from being cancelled prematurely.How often should I review search performance?
Weekly for faults only — Manual Actions, Security Issues, Crawl Stats host status, indexed count direction, server errors — which should take ten minutes and generate no discussion in a normal week. Monthly for performance by segment against the baseline, plus what shipped and what is blocked. Quarterly for the direction decision. Avoid reporting on cycles shorter than the thing being measured: field performance data accumulates from real visits, indexing responds to recrawls, and weekly reporting on either is an invitation to react to noise.Why should I separate branded from non-branded search in reports?
Because branded search mostly measures your marketing everywhere else. When someone searches your company name, they already know you, and the search channel is capturing demand rather than creating it. Leaving branded terms in a search report lets a television campaign, a funding announcement or a product launch appear as an optimization result — and it works in reverse too, making a genuine improvement invisible underneath a decline in brand demand. Split them in every report. If only one segmentation is possible, this is the one to keep.How do I tell whether a change was caused by a Google update?
Test the dates rather than assume. Google publishes a ranking release history with start dates and durations. The movement should begin when the rollout began and stabilize around when it completed. Precision matters: the March 2026 core update completed in under twenty hours and a separate spam update began three days later, so a late-March change involved two distinct events that only dated, segmented data can separate. If your movement started weeks before any published window, the update is not the explanation and the investigation should continue.Should my report include a single overall SEO score?
No. A composite score blends measures with different meanings, different reliability and different blind spots into one number that cannot be acted on, and it is ultimately the opinion of whichever tool produced it. Two vendors will produce different scores for the same site because they are weighting different inputs. Report the measures separately, each with its window and segment stated, and add one short paragraph on what the data cannot see. A decision-maker can act on that; they cannot act on a number out of a hundred.What should a report say about things it cannot measure?
Say them specifically, every time, in a short paragraph rather than a disclaimer. Useful examples: which segments fell below Search Console's 1,000-row export cap, where consent controls or tracking prevention leave sessions uncounted, that rare queries are withheld so filtered totals will not sum to unfiltered ones, and that AI feature appearances are reported without any click data. Stating limits in month one converts them from excuses in month nine into shared understanding, and it makes the numbers you do report considerably more credible.Published