I keep seeing the same question in scraping threads, usually from someone whose side project just got promoted into a real internal tool: how do I actually fetch search results at scale, and what should I be paying for? The honest answer is that there are three different things you can buy — a residential proxy pool, a hosted scraping browser, or a managed SERP API — and which one wins depends almost entirely on one number nobody bothers to measure: how many bytes is one search results page, really?
So I measured it, then I built a small budget model around the number, and I want to walk through what it says. This is not a neutral benchmark; I work with Thordata, so treat it as a worked example on infrastructure I help sell. But the arithmetic below is copy-paste runnable, and it does not flatter the more expensive options. If the numbers say buy the cheap pool, they say that.
What a single SERP page actually weighs
I pulled five product-intent queries through a plain HTTP GET (with --compressed) against a public search endpoint and logged only the transfer size and wall time. No proxy, no browser, no cleverness — this is the raw payload your pipeline has to move and then parse, per request.
UA="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
for q in "running+shoes+price" "budget+laptop+deals" "espresso+machine+compare" \
"air+fryer+best+price" "hiking+boots+sale"; do
curl -sL --compressed -A "$UA" -o /dev/null \
-w "gzip: code=%{http_code} wire=%{size_download} time=%{time_total}\n" \
"https://www.bing.com/search?q=$q&count=20"
curl -sL -A "$UA" -o /dev/null \
-w "plain: wire=%{size_download} time=%{time_total}\n" \
"https://www.bing.com/search?q=$q&count=20"
sleep 1
done
Median of five, taken on 2026-09-28:
-
~99 KB of wire bytes with a naive, uncompressed client (
97889..100074) -
~35 KB with gzip on (
34365..35689) - ~0.5 s per round trip
That 3× gzip delta is the whole ballgame, because two of the three billing models are per-GB. Whether you remembered Accept-Encoding: gzip is, to a bandwidth-metered product, a direct 3× swing in cost. Hold that thought.
The multiplication that turns one page into a bill
A rank tracker is not one request. It's keywords × tracked locations × refreshes per day × device passes. A mid-size commerce account — 500 keywords, 12 metro areas, checked four times a day, desktop and mobile — is not 500 fetches. It's:
500 x 12 x 4 x 2 x 30 = 1,440,000 responses / month
Times ~35 KB each (gzip) and you are moving about 51 GB per month. Times ~99 KB (if you forgot gzip) and you are at ~143 GB. Same data, same job. The only difference is one header, and it's a ~2.8× line item.
I want to be able to replay this without opening a spreadsheet, so I wrote the whole model as a bash + awk script. No runtime, no packages — bash and awk are on any box you're already scraping from. Live prices are pulled from the Thordata pricing page the same morning; re-run the curl in the comment before you trust any dollar figure, because promos move.
#!/usr/bin/env bash
# serp_line_budget.sh - size a rank-tracking job against three collection lines.
# Live prices pulled from thordata.com/pricing/serpapi on 2026-09-28.
# Re-verify before publishing: curl -s https://www.thordata.com/pricing/serpapi
# Usage: bash serp_line_budget.sh <keywords> <locations> <runs_per_day> [device_passes]
set -euo pipefail
KW=${1:?keywords}; LOC=${2:?locations}; RPD=${3:?runs/day}; DEV=${4:-2}
# --- measured 2026-09-28, plain GET to a SERP page, median of 5 samples ---
BYTES_RAW=99230 # range 97889..100074
BYTES_GZ=35338 # range 34365..35689 (curl --compressed)
# --- live unit prices (pick the tier your monthly volume lands in) --------
RESI_GB=0.80 # residential, 350GB tier; high-volume down to $0.65
SB_GB=2.50 # scraping browser, 500GB tier ($5 at 1GB, $2.50 at 500GB)
SERP_1K=0.80 # SERP API, 500k-response tier ($0.70 at 1M, $1.20 at 15k)
awk -v kw="$KW" -v loc="$LOC" -v rpd="$RPD" -v dev="$DEV" \
-v br="$BYTES_RAW" -v bg="$BYTES_GZ" \
-v rg="$RESI_GB" -v sbg="$SB_GB" -v s1k="$SERP_1K" '
BEGIN{
req = kw*loc*rpd*dev*30 # responses/month (desktop+mobile passes)
e_r = req*br/1e9 # egress GB, raw HTML
e_g = req*bg/1e9 # egress GB, gzip
res = e_r*rg # residential pool, $/mo (you still parse)
brows= e_g*sbg # scraping browser, $/mo (renders + you parse)
api = (req/1000)*s1k # SERP API, $/mo (clean JSON, no parsing)
eng = 9000 # loaded $/mo of 0.5 engineer to own the fleet
printf "Job: %s kw x %s cities x %s/day x %s passes = %d responses/mo\n\n", kw,loc,rpd,dev,req
printf "Egress %6.1f GB raw | %6.1f GB gzipped\n", e_r, e_g
printf "Residential pool $%7.0f/mo (+ you own the HTML parser, retries, fleet)\n", res
printf "Scraping browser $%7.0f/mo (+ JS render handled; you still parse)\n", brows
printf "SERP API $%7.0f/mo (JSON out; nothing to babysit)\n", api
printf "\nIf the DIY path needs half an engineer ($%d/mo), API beats DIY when\n", eng
printf "req*(price gap) < %d. Break-even responses/mo ~= %d.\n", \
eng, eng/ (s1k/1000 - rg*bg/1e9)
}'
Run it on the mid-size account:
$ bash serp_line_budget.sh 500 12 4 2
Job: 500 kw x 12 cities x 4/day x 2 passes = 1440000 responses/mo
Egress 142.9 GB raw | 50.9 GB gzipped
Residential pool $ 114/mo (+ you own the HTML parser, retries, fleet)
Scraping browser $ 127/mo (+ JS render handled; you still parse)
SERP API $ 1152/mo (JSON out; nothing to babysit)
If the DIY path needs half an engineer ($9000/mo), API beats DIY when
req*(price gap) < 9000. Break-even responses/mo ~= 11662115.
Read that honestly: at the raw-dollar level, the residential pool wins by an order of magnitude. 1.44M responses for $114 versus $1,152 for the API. Anyone who tells you the managed API is "cheaper" is selling you something else. What the API costs ten times more than is not the bytes. It's the parsing.
So what are you actually buying?
Three different products, three different amounts of work left on your desk after the fetch:
Residential pool. You get real IP diversity and the cheapest per-GB egress. In exchange you own the HTML: the selectors, the layout-change monitoring, the "why did 8% of this run come back empty" debugging, the connection pool, the retries, the geo-anchoring, the whole fleet. The $114 line in that table is the transport bill. It is not the total cost of ownership.
Scraping browser. Same per-GB shape, but the provider runs a real headless browser and hands you the rendered DOM. The line item barely moves ($127 vs $114) because the extra money buys you JavaScript execution, not structured data. If a page is a server-rendered soup, the residential pool is enough and this is a slight premium for nothing. If the page is a JS app that hydrates results client-side, the browser earns its 13-cent premium instantly — parsing an empty pre-hydration shell is the worst money you can spend.
SERP API. The one that costs ten times more, and the only one that returns clean JSON instead of HTML. This is where the "what are you buying" question has a real answer: you're buying your way out of the parser. A SERP API already knows that a blue-link result lives at organic[].link, that ads are ads[], that "position zero" is a featured snippet — and it relearns that mapping every time the search engine redesigns, on their clock, not yours.
Now put the $9,000 engineer line back in. If owning the fleet genuinely costs you half a person a month, the break-even printed above says the API is the cheaper total choice up to roughly 11.6M responses/mo. Your mid-size account at 1.4M is nowhere near that ceiling — which is the honest way to say it: for most small and mid teams, a managed SERP API pays for itself against engineer time, even though it costs 10× the raw bandwidth. The bandwidth was never the expensive part. The keeping-warm of a parser was.
A worked pick, not a philosophy
If I were wiring this today, I'd map the shape of the job to the line:
-
You're tracking Google/Bing SERPs as data — ranks, featured snippets, ad presence — into a table or a warehouse, at daily-or-slower cadence. Buy the SERP API. You want JSON, you want the parse handled, and your egress bill was never going to be the constraint. Thordata meters it per successful response (
/1K), with a free 500-response tier to wire up first, and a 7-day free trial. - You're scraping JS-heavy result pages that aren't classic SERPs — marketplace search, job boards that hydrate client-side, price pages behind a SPA. Buy the Scraping Browser, parse once, keep the parser under version control. Your cost is dominated by egress, so keep gzip on and measure per-object size like in my throughput pieces.
- You already have a working parser and just need IP diversity and geo anchoring at the lowest per-GB rate. Buy the residential pool and keep the browser work local. This is the right call precisely when parsing is already your core competence — the table shows it's the cheapest line by a mile.
The failure mode I see in practice is teams defaulting to the residential pool because it's cheapest per byte and then quietly spending more on engineering than they'd ever have paid the API. If the SERP page is the product you're shipping — you're building SEO tooling, you're benchmarking rankings — buy the API. If the SERP page is an incidental fetch on the way to something else, buy the pool. The measured ~35–99 KB per page is just the input that lets you do that arithmetic on your own volume instead of guessing.
Two limitations I'll own rather than bury: the SERP page sizes above are one engine, five samples, product-intent queries — informational and navigational queries weigh differently, and image/video result pages weigh a lot more. And the live prices are one morning's snapshot against a public tier ladder; the residential per-GB rate slides from $2.00 at 1 GB down to $0.65 at 5000 GB, so re-run the curl and drop in your own tier before you quote any number from this post.
Disclosure: I work with Thordata. The budget script here is generic and MIT-spirited — swap in any provider's unit prices and it'll do the same honest arithmetic, even when that arithmetic tells you to buy the cheap pool.













