Our Next.js pages got too large for Bing. RSC payloads were 58% of the HTML
Three Next.js pages stalled at discovered but not crawled in Bing. The App Router ships every byte twice. We cut one by 48 KB and Bing indexed it.
October 7, 2026
Katto is an AI video clipper. You hand it a long video, it finds the best moments and cuts them into captioned vertical clips. I build it on my own, in public, on Next.js with the App Router. This is a measurement that changed how I write pages, a number I got wrong in front of everyone, and one hypothesis of mine that turned out to be worth 108 bytes.
The symptom
Bing Webmaster Tools listed several of our pages as discovered but never crawled. Not blocked by robots.txt, not noindex, not slow, not redirecting. Discovered, then dropped. Google had them indexed. Our sitemap listed them. They returned 200 to anything that asked.
The pages it refused had one thing in common, and it was not their content.
The boundary, and how loosely we know it
Microsoft does not publish a size at which its crawler gives up. What we have is our own pages, each one with its status read in Bing Webmaster Tools and its served size measured the same day, uncompressed, with bingbot's user agent:
/contact, 46,111 bytes: indexed./changelog, 103,058 bytes: indexed./roadmap, 116,239 bytes: indexed./compare/opusclip, 127,182 bytes: indexed./compare/ai-video-clippers, 127,951 bytes: indexed./compare/, 140,686 bytes: discovered but not crawled.
One page on the far side of a line is not a law. But there is a single observation in that list worth more than the rest of it, because it is the same URL twice. /compare/ai-video-clippers served 172,295 bytes and sat at discovered but not crawled for weeks. We cut it to roughly 124 KB. It is indexed today, at 127,951 bytes. Same URL, same template, same kind of content: too large, then small enough.
A week ago I would have given you a different number, and it would have been wrong. I had three status readings and a conspicuous gap in our size distribution, and I read a boundary near 125 KB into it. Two of those pages have since been indexed at just under 128 KB. What our own site can actually say is that the line sits somewhere between 128 and 140 KB, and that the far side of it is still exactly one page wide.
If you came here for a number, the honest answer is to measure your own. The claim I can defend is narrower and more useful anyway: a page that stalls at discovered is worth weighing before it is worth rewriting, and a page that starts being crawled after you shrink it is the only evidence that really counts.
Weighing the source file tells you nothing
Here is the part that makes this specific to React Server Components. The App Router serialises the rendered tree into the same HTML document, as a run of self.__next_f.push(...) script tags, so a client-side navigation can resume without another round trip. That payload is not a separate request a crawler can skip. It is in the document.
Measured on one of our pages: 125,981 bytes served, of which 73,571 are that payload. Fifty eight percent of the document is the serialised copy of what the other forty two percent already says.
So a utility class written once inside a loop over 24 rows does not ship 24 times. It ships 48.
The part that took me longest to see
It does not double the page. It doubles every representation the page already carries.
Our roadmap deliberately rendered each entry twice: once as visible text, once inside an ItemList block of JSON-LD, so machine readers would get the name, the date and the status as data rather than prose to interpret. Reasonable. Then I counted the occurrences of a single entry's description in the served document and got four.
Two representations, each doubled. Four copies of the same sentence, and I had never thought to count.
The fix was not compression. The JSON-LD carried a description that was a verbatim copy of text the reader sees thirty lines below it, in the same language. It bought nothing and cost two of the four copies. Removed, the block keeps the name, the date and the status, which is what prose does not state explicitly to a machine. That page went from 121,196 to 115,350 bytes and no information left it. Both figures come from the same build, before and after that one change, which is the only way a before and after means anything.
If you take one habit from this: before trying to shrink anything, count how many times one sentence of your content appears in the served document. It answers a question that weighing sections does not.
The two bigger levers
Repeated class attributes into stylesheet rules. Tailwind utilities are wonderfully cheap to write and they ship on every element, twice. Moving the repeated ones into @apply rules took our pricing index from 172,295 bytes to roughly 142,000. They were about 51 percent of the markup.
Non-interactive blocks out of the serialised tree. To be precise, because it matters: these are Server Components, so they are never hydrated on the client. The cost is not hydration, it is serialisation. Every element, prop and key is written into the flight payload so a client-side navigation can rebuild the tree without a round trip. A static table with no client behaviour still pays that, and when document size is the constraint there is no reason it should remain React elements at all. We now build those blocks as an HTML string from the same data and set them with dangerouslySetInnerHTML, escaping every interpolated value at construction so the escaping does not depend on where the data came from. Another 18 KB on the pricing index. A 54 row list on the roadmap went from 157,308 to 119,031 the same way.
The thing that did not work
Before any of that, I was sure the component boundaries were the problem. Fifty four next/link components in a list, each contributing its serialised name, props and keys to the payload. I replaced all 54 with plain anchors.
It saved 108 bytes out of 157,200.
Link is cheap. The content it wraps is not. I would have spent a day rewriting components on that theory if I had not measured before and after, and the negative result was the most useful number of the day.
There is a floor, and it is not small
Our lightest real page, /contact, serves 46,111 bytes while carrying a heading, a short form and a few questions. The rest is navigation, footer, analytics and fonts. Below that figure there is nothing to win page by page; it would be a shell project, not a page project. Knowing the floor tells you when to stop.
A translation costs bytes, not words
We put that roadmap into ten languages. Same markup, same structure, same information. The English page serves 116,239 bytes, French 119,838, Japanese 122,155 and Hindi 135,379.
Devanagari takes three bytes per character in UTF-8 where Latin text takes one, and the payload pays for it a second time. No page level technique touches that: same markup, same information, nineteen kilobytes apart.
The Hindi page is the interesting one, because at 135,379 bytes it sits inside the band we cannot resolve. Whether Bing processes it is the cheapest experiment available to us, and it costs one URL inspection.
I had written a paragraph here arguing that the overflow did not matter much, because it falls at the end of the document and the end of that page is a list whose titles stay in English anyway. I cut it, because it assumes the crawler reads until it runs out of budget and then stops. Everything I have actually observed says only that pages above a certain size are not processed. If the document is rejected whole, its ordering buys nothing, and I have no evidence either way. It is a comfortable theory about my own page, which is the kind most worth deleting.
It is a budget, not a fix
The pricing index I took from 172,295 to 124,202 bytes is at 127,951 today, because we kept adding to it: more tools, more columns, ten translations. It is still indexed, which is the whole point, and it is also back within a few kilobytes of the line. Nothing went wrong. We simply spent the budget again without watching it.
That is the real conclusion. Page weight in a framework that serialises its own output is not a cleanup you perform once. It is a number that drifts up every time someone does good work, and the only defence is to measure the served document rather than the source, one change at a time.
How to measure yours
- Request the URL and count the bytes of the response body, uncompressed. That is what the crawler receives, and it is not what your source file suggests.
- Count the occurrences of one distinctive sentence of your content in that body. More than one means you are paying for it more than once, and it points at the duplicate.
- Re-measure after every single change, separately. Two changes at once hide the one that did nothing.
- Then read the status in Bing Webmaster Tools again, on the same URL. A boundary you infer from other people's pages is a guess. A page of yours that starts being crawled after you shrink it is a measurement.
Related articles
Ready to turn your videos into viral clips?
Katto automatically clips, captions, and reframes your long-form videos into short-form content.
Try Katto for free →