She had not changed anything. The graph moved anyway.
Leila had done everything the right way. She had warmed her sending domain slowly — fifty emails per day in week one, scaling carefully over eight weeks before reaching full volume. She had sourced her lists from reputable providers and verified them before importing. She had kept her sequences tight, her ICPs specific, her sending volume consistent. For eleven months, none of this had produced a deliverability problem worth mentioning.
Then in March, she opened Google's Postmaster Tools on a Tuesday morning and the domain reputation graph had moved from high to medium. By Thursday it was low. By the following Monday her open rates had dropped from 22 percent to 4 percent.
What I found when I dug into her sending domain: the drop had not been caused by one thing. It had been caused by four things, across six days, none of which she had directly produced, none of which she had been able to see while they were happening. She had not failed at outbound. She had failed at monitoring. Those are different problems, and they have different fixes.
Why the graph moved on Thursday when the problems started on Saturday
Domain reputation in Gmail is a classification, not a score. The four levels reflect how Gmail's model has assessed the aggregate sending behavior of a domain across a rolling window. The model does not update in real time and it does not react to single events in isolation. It accumulates signal.
Most teams, when they see the classification drop on Tuesday, spend Tuesday looking for something that went wrong on Monday. When they cannot find anything on Monday, they conclude Google's model must be wrong. The real answer is almost always in the days before Monday. The cause that moved Tuesday's classification was already fully in motion by the previous Wednesday.
When I worked through Leila's situation, I went back ten days from the classification change and rebuilt what had happened to her sending domain, day by day. What I found was a sequence that started on February 28th and produced a visible consequence on March 4th. Four events. Each small. Together, threshold-crossing.
Four events. Six days. One classification change.
A sales partner sent Leila's team a list of contacts from a joint account database. The partner ran a legitimate operation. The list had been compiled carefully. Nobody had anything to gain from sending Leila bad data. The list was just old.
B2B contact data loses accuracy at roughly 2 to 3 percent per month. Over fourteen months, a list of 1,800 contacts accumulates somewhere in the range of 350 to 680 addresses that are no longer accurate. The verification audit run after the fact came back with 391 hard invalid addresses.
A VP of Operations had been in Leila's sequence for three weeks. He had opened the first two emails and then stopped engaging entirely. By touchpoint four he had not opened or clicked anything in seventeen days. The email arrived. He marked it as spam.
One complaint. Leila had no visibility into it. Gmail does not notify senders of individual spam complaints. But the complaint did not arrive in isolation — it arrived twenty-four hours after the bounce event, on a domain already accumulating negative signal. Gmail's classification model does not evaluate events in isolation. They compound.
An automated spoofing kit had added her domain to its rotation and began sending phishing emails claiming to originate from her domain. The volume was modest — 80 to 120 messages per day. Most were caught by Gmail's filters. But some got through, and recipients who received those messages and marked them as spam generated complaint signals attributed to her domain.
Her DMARC reporting mailbox had been receiving daily aggregate reports showing every IP sending from her domain. The reports had been arriving for eleven months. She had never opened the mailbox. The spoofing had started March 3rd. It was in the report that arrived March 4th — the same day the classification changed.
When 391 addresses in a sequence receive an email and bounce, they count as sent messages against which zero engagement is recorded. Leila's domain had been running consistent open rates of 18 to 22 percent. The 391-contact bounce cohort injected a zero-engagement block into the same rolling window the domain reputation model uses to assess whether recipients want mail from this sender.
The diagnostic sequence — ordered because earlier events explain later ones
Every one of those steps is a monitoring practice that should have been running continuously. The diagnostic process is what it looks like when all five checks happen reactively — after the drop — rather than proactively. Running them proactively would have caught each of the four events individually, before they had a chance to compound.
Seven weeks. Not permanent — but not fast either.
Leila ran all five steps in the week following the call. The 391 invalid addresses were suppressed. Engagement exit conditions went into every active sequence. The DMARC policy moved from p=none to p=quarantine. Every other list in her CRM was audited for age and re-verified before any further sends.
A team that identifies the cause late, fixes only some of the contributing factors, or continues sending at full volume to unfiltered contact lists during recovery will take longer — sometimes significantly longer. Seven weeks was the result of doing it right after the diagnosis. Leila's eleven months of careful background work was the reason recovery was possible in seven weeks rather than seven months.
Three practices. None requiring a new tool or a technical hire.
None of these require a new tool. None require a technical hire or a budget allocation. They require calendar reminders and decision rules written into sequence setup. They are the difference between a team that catches individual events before they compound and a team that discovers what compounded after the graph moves.
Two things being built simultaneously — only one of them visible
She said she had not understood, until the graph dropped, that she had been building two things simultaneously for eleven months. She had been building pipeline through her outbound sequences. And she had been building domain reputation through every list decision, every verification choice, every send timing call, every engagement exit she did or did not make.
The pipeline results had been visible in her CRM every week. The domain reputation had been invisible — right up until it was not.
The gap between a domain that holds at high reputation across a year of consistent outbound and a domain that drops in six days is almost always in those background decisions — not in the sequence quality, not in the ICP, not in any of the foreground variables that get most of the attention.
Accurate results at the point before the damage is done
Leila's 391-bounce event began with a partner list that never went through verification before import. Two hours and a verification run would have caught every address in it. That is what Proxy25 is built for.