The Web Scraping Club

298 posts

The Web Scraping Club

The Web Scraping Club

@webscrapingclub

The Web Scraping Club is a substack with news, tutorials, real-world code examples, and anti-bot bypass examples.

Milan Katılım Şubat 2023
176 Takip Edilen705 Takipçiler
Sabitlenmiş Tweet
The Web Scraping Club
The Web Scraping Club@webscrapingclub·
Google has deindexed my entire Substack substack.thewebscraping.club. 4 years of work. Hundreds of posts. Gone from the search overnight. The nightmare every content creator fears just happened to me.
The Web Scraping Club tweet media
English
2
1
3
225
The Web Scraping Club
The Web Scraping Club@webscrapingclub·
If one of my posts has ever helped you, the best thing you can do right now is share it with ONE person who works in scraping, automation, or data engineering. @Substack @Google, any help appreciated.
English
1
0
0
75
The Web Scraping Club
The Web Scraping Club@webscrapingclub·
But none of that matters when Google decides you no longer exist. In the meantime, I'm back to relying on the oldest growth channel: word of mouth. It's been my main driver until now. From here on, it's everything.
English
1
0
2
77
The Web Scraping Club
The Web Scraping Club@webscrapingclub·
Google has deindexed my entire Substack substack.thewebscraping.club. 4 years of work. Hundreds of posts. Gone from the search overnight. The nightmare every content creator fears just happened to me.
The Web Scraping Club tweet media
English
2
1
3
225
The Web Scraping Club
The Web Scraping Club@webscrapingclub·
Most e-commerce scrapers waste bandwidth fetching unchanged product pages repeatedly. We tested HTTP conditional requests against Shopify stores and found 95% bandwidth savings using ETags and 304 responses. One simple header can cut proxy costs dramatically. substack.thewebscraping.club/p/http-caching…
English
1
0
4
121
The Web Scraping Club
The Web Scraping Club@webscrapingclub·
This isn't just about SerpApi. If Google's theory prevails, every CAPTCHA and bot detection system on websites with copyrighted content could invoke federal law against scrapers. The entire industry's legal framework is at stake.
English
1
0
0
84
The Web Scraping Club
The Web Scraping Club@webscrapingclub·
Google filed a 12-page lawsuit against SerpApi using the DMCA anti-circumvention provision - the same law designed to stop DVD piracy. No cease-and-desist. No communication. Straight to federal court.
English
1
0
2
130
The Web Scraping Club retweetledi
Apify
Apify@apify·
66% more proxy usage. 86% report increased anti-bot protections. Only 46% using AI. Apify + @webscrapingclub surveyed scraping pros on what's working in 2026. The full report covers proxies, infrastructure, bot detection, and AI trends. Link in thread 👇
Apify tweet media
English
2
2
9
679
The Web Scraping Club
The Web Scraping Club@webscrapingclub·
Selenium, Playwright, Puppeteer. You use them daily. But do you know how they actually control browsers? Three protocols: - WebDriver: HTTP, cross-browser, polling - CDP: WebSocket, Chromium-only, real-time - BiDi: WebSocket, cross-browser, real-time BiDi is the future. It won't replace CDP. Link: substack.thewebscraping.club/p/webdriver-vs…
The Web Scraping Club tweet media
English
0
0
0
88
The Web Scraping Club
The Web Scraping Club@webscrapingclub·
We scraped 1000 #Nike pages with 5 tools. HTTP client: 8 minutes Browser automation: 3.5 hours Same data. Same success rate. 27x speed difference. Nike checks TLS fingerprints, not JavaScript execution, at least in the product catalog. The data is in the first response. Stop paying the browser tax when you don't need to. Link: substack.thewebscraping.club/p/scraping-nik… #webscraping
The Web Scraping Club tweet media
English
0
0
1
76
The Web Scraping Club
The Web Scraping Club@webscrapingclub·
Have you ever wondered how #MachineLearning can be used to detect #bot traffic? In the latest post of The Web Scraping Club, Federico Trotta showed us some details about it. From unsupervised training to supervised one, traffic data can be clustered by user behavior and system settings, identifying some patterns that distinguish bots and human. Link to article: substack.thewebscraping.club/p/machine-lear…
English
0
0
2
341