Seventeen posts disappeared from the abd.dev sitemap.
The posts still existed. Their pages still worked. The homepage could still link to them.
But Google,
llms.txt and anything else reading the sitemap were being shown an incomplete site.The strange part was that I had already added retry logic to make the Notion crawl more reliable.
That retry logic was causing the failure.
It worked locally
abd.dev uses Notion as its CMS. During a build, the site crawls the Notion workspace and generates the pages, sitemap and other indexes from what it finds.
A local build found everything.
The Vercel build did not.
That difference mattered. Vercel runs multiple workers during static generation, and each worker can trigger more calls into Notion. Enough concurrent requests eventually hit Notion's rate limit.
Rate limiting by itself should have been recoverable. The Notion client already knew how to recover from it.
My configuration prevented it from doing so.
The helpful override
The client had been configured like this:
const notion = new NotionAPI({ ofetchOptions: { retry: 5, retryDelay: 2000, }, })
Five retries. Two seconds between them.
That looks reasonable if you think every failure is a brief network problem. Try again quickly, recover and continue the build.
But a rate limit is not a random failure. It is the server telling you when trying again can become useful.
The Notion client already had specific handling for HTTP 429 responses. It waited for the rate-limit cooldown, which was usually around 60 to 75 seconds.
My fixed
retryDelay overrode that behavior.Every retry happened after two seconds. Every retry landed inside the same cooldown. The client exhausted all five attempts before Notion was ready to accept another request.
The crawl then continued without those pages.
The deployment succeeded. The sitemap was incomplete.
The dangerous part was not the error
There was no dramatic outage.
The build stayed green. The site loaded. Existing links worked.
The failure degraded the completeness of the output instead of stopping production entirely.
That is harder to notice because all the obvious health checks still pass. A request to the homepage returns 200. A request to the sitemap also returns 200. Neither tells you whether the sitemap contains everything it should.
We even had a warning intended to report skipped pages. It could never fire. Failed pages were dropped inside the library before our code received the crawl result.
We were monitoring a failure state our layer could no longer observe.
The fix was deleting code
I removed the retry override and let the Notion client own the cooldown.
const notion = new NotionAPI()
Then I tested both configurations against a fake Notion server that returned a 429 once and succeeded afterward.
The old configuration retried after 2.0 seconds.
The corrected configuration retried after 66.6 seconds.
The longer delay looked worse if the only metric was retry speed. It was the only delay that actually recovered.
After deployment, the missing posts returned to the production indexes. We also added a nightly check that compares the site's visible posts with the sitemap so an incomplete crawl cannot quietly pass again.
Retry logic needs to understand the failure
Retries are often treated as generic resilience:
Something failed. Wait a little. Try again.
But different failures require different recovery behavior.
A connection reset may justify an immediate retry. A rate limit requires waiting for the server's cooldown. An invalid request should not be retried at all. A partially completed write may require checking what happened before repeating anything.
Flatten all of those into one fixed interval and the retry loop can amplify the original problem.
The rule I carry now is simple:
If a client library understands the protocol better than my wrapper does, let it own the recovery behavior.
Do not replace specific backoff logic with a generic retry because the generic version looks simpler.
Sometimes the fastest way to recover is to stop retrying so fast.
