The Web Scraping Club
Subscribe
Sign in
Home
News
The Lab
Advertise on TWSC
Proxy pricing benchmark
Consulting
Archive
About
The Lab
Latest
Top
Discussions
THE LAB #103: Bypassing DataDome-Protected Websites in the Agentic Era
Fifteen browser configurations, one tough anti-bot, and only a couple made it to the cart
May 1
•
Pierluigi Vinciguerra
4
THE LAB #102: How Fast Can You Call Polymarket's APIs?
Three languages, four locations, 1,000 requests. The biggest speed gain has nothing to do with code.
Apr 16
•
Pierluigi Vinciguerra
4
THE LAB #101: Building an Internal Knowledge Base for Your Scraping Team
Every scraping team that survives long enough develops the same disease.
Apr 3
•
Pierluigi Vinciguerra
1
1
THE LAB #100: Hybrid Scraping - One Browser Login, Thousands of HTTP Requests
Building a pipeline that uses Camoufox for authentication and curl_cffi for extraction on Akamai-protected targets.
Mar 20
•
Pierluigi Vinciguerra
3
1
THE LAB #99: HTTP Caching for Web Scraping
How Conditional Requests Can Cut Your Proxy Bill, using HTTP caching.
Mar 6
•
Pierluigi Vinciguerra
1
1
THE LAB #98: Scraping Google Search Results in 2026: Device, Location, and Identity
Google does not have one set of results. It has millions. The hard part is knowing which one you are looking at.
Feb 19
•
Pierluigi Vinciguerra
1
2
1
THE LAB #97: My first week with OpenClaw
160,000 Stars in Two Months: What OpenClaw Means for Scrapers
Feb 6
•
Pierluigi Vinciguerra
6
2
THE LAB #96: Scraping Nike.com with 5 open source tools
Match your tool to the protection, not the brand
Jan 30
•
Pierluigi Vinciguerra
6
THE LAB #95: Bypassing Cloudflare in 2026
Testing Open Source Browser Automation Tools Against Real Targets
Jan 23
•
Pierluigi Vinciguerra
3
THE LAB #94: Using cookies and session for cost-effective scraping
Trying to scrape from Leboncoin, Allegro and Idealista without emptying our pockets.
Oct 18, 2025
•
Pierluigi Vinciguerra
6
1
THE LAB #93: scraping Booking.com using internal APIs
How to get travel data without a browser
Sep 20, 2025
•
Pierluigi Vinciguerra
2
THE LAB #92: scraping Depop in a cost-effective way
Delegate or in house solution? A cost-wise perspective.
Aug 30, 2025
•
Pierluigi Vinciguerra
5
THE LAB #91: Performing sentiment analysis on Amazon product reviews - Part #2
A Practical Guide to Sentiment Analysis on Data Scraped From Amazon
Aug 14, 2025
•
Federico Trotta
3
1
THE LAB #90: Camoufox Server in AWS
How to create a cluster of Camoufox instances on AWS
Aug 1, 2025
•
Pierluigi Vinciguerra
3
4
THE LAB #89: Camoufox as a Docker image
Step by step guide to scale your Camoufox web scraping infrastructure
Jul 17, 2025
•
Pierluigi Vinciguerra
3
6
1
THE LAB #88: Fuel Your Content Machine with LLM Scraping
Fill your Obsidian Vault with daily inspiration from the web
Jul 4, 2025
•
Pierluigi Vinciguerra
4
1
THE LAB #87: Bypassing ReCAPTCHAs with open source and commercial tools
History, Technical details, Alternatives, and Bypass Methods of ReCAPTCHA
Jun 21, 2025
•
Pierluigi Vinciguerra
3
2
1
THE LAB #86: Querying Web Data using GPT-Like Web Interface
How LLM disrupted the concept of self-service business intelligence
Jun 5, 2025
•
Pierluigi Vinciguerra
5
THE LAB #85: Bypass Akamai Bot Protection by Chaining Proxies
How JA3Proxy and residential proxies can help us bypassing Akamai
May 29, 2025
•
Pierluigi Vinciguerra
8
1
3
THE LAB #84: AI-Driven Web Scraping: OpenAI Codex vs Cursor vs AI Scraping Tools
Is OpenAI Codex the new silver bullet for scraping?
May 22, 2025
•
Pierluigi Vinciguerra
4
THE LAB #83: Camoufox as a containerized server
Upgrade your scraping tech stack with Camoufox server and scale its deployment
May 16, 2025
•
Pierluigi Vinciguerra
8
2
1
THE LAB #82: How to scrape Vinted using its internal APIs
Why analyzing both app and websites is key for a successful web scraping project
May 1, 2025
•
Pierluigi Vinciguerra
10
1
2
THE LAB #81: Scraping Zillow for fun and profit
Get real estate data from the biggest US website
Apr 17, 2025
•
Pierluigi Vinciguerra
4
1
THE LAB #80: Scraping food delivery data
Use both the website and the mobile apps to get data from food and grocery delivery data
Apr 3, 2025
•
Pierluigi Vinciguerra
1
THE LAB #79: Use Cursor as a web scraping assistant with MCP servers
Add MCP Servers to Cursor for increasing our web scraping capabilities
Mar 21, 2025
•
Pierluigi Vinciguerra
12
THE LAB #78: Building a Web Scraping Knowledge Assistant with RAG - Part2
Optimizing the content storage and creating a CLI for our assistant
Mar 7, 2025
•
Pierluigi Vinciguerra
6
THE LAB #77: Building a Web Scraping Knowledge Assistant with RAG
How to include scraped data in your AI assistant
Feb 28, 2025
•
Pierluigi Vinciguerra
13
THE LAB #76: Bypassing Kasada With Open Source Tools In 2025
How to bypass Kasada protected websites without paying a cent
Feb 14, 2025
•
Pierluigi Vinciguerra
6
THE LAB #75: Building self healing scrapers with AI
How can we use LLMs to analyze HTML and fix our web scrapers?
Feb 6, 2025
•
Pierluigi Vinciguerra
8
THE LAB #74: Running scrapers on GitHub Actions
Save money and time by using GitHub infrastructure for running your scrapers
Jan 30, 2025
•
Pierluigi Vinciguerra
8
THE LAB #73: How to Bypass Cloudflare in 2025
Scraping websites protected by Cloudflare bot protection with open source tools
Jan 23, 2025
•
Pierluigi Vinciguerra
16
4
1
THE LAB #72: Advanced logging in Playwright
RabbitMQ, screenshots and system monitoring for your Playwright scrapers
Jan 10, 2025
•
Pierluigi Vinciguerra
6
1
THE LAB #71: Sending Scrapy logs to RabbitMQ
Saving the logs of your distributed scraping architecture to your database
Dec 19, 2024
•
Pierluigi Vinciguerra
1
THE LAB #70: Advanced logging in Scrapy
How to extract the most meaningful metrics from your scrapers
Dec 13, 2024
•
Pierluigi Vinciguerra
5
THE LAB #69: Building a dashboard for your scrapers with Grafana
Visualizing the operations of your Scrapy spider with interactive dashboards
Dec 5, 2024
•
Pierluigi Vinciguerra
8
THE LAB #68: Scheduling Scrapers with Airflow
How to manage a fleet of scrapers with Apache Airflow
Nov 28, 2024
•
Pierluigi Vinciguerra
5
1
THE LAB #67: Scraping Telegram using its APIs
How to create a bot for scraping Telegram channels
Nov 21, 2024
•
Pierluigi Vinciguerra
11
THE LAB #66: How to properly scrape a booking website
Business logic and best practices for scraping booking websites like Airbnb and Booking in the most efficient way
Nov 8, 2024
•
Pierluigi Vinciguerra
3
THE LAB #65: Scraping Datadome protected websites with Camoufox
Discovering the features of Camoufox, a custom and stealthy version of Firefox
Oct 24, 2024
•
Pierluigi Vinciguerra
8
2
THE LAB #64: JWT Tokens and API scraping
How to create scrapers that use token authentication for API data retrieval
Oct 17, 2024
•
Pierluigi Vinciguerra
3
1
THE LAB #63: Oxymouse and Playwright for human-like mouse movements
Testing the new Oxylabs open source package for human-like mouse movements
Oct 4, 2024
•
Pierluigi Vinciguerra
3
THE LAB #62: Bypassing Cloudflare with Nodriver
Testing the undetected-chromedriver successor for scraping Cloudflare protected websites
Sep 26, 2024
•
Pierluigi Vinciguerra
2
THE LAB #61: Evaluating your proxy provider
Measuring programmatically the quality of the IPs offered by proxy providers
Sep 12, 2024
•
Pierluigi Vinciguerra
9
THE LAB #60: Writing scrapers with LLMs
Comparing LLama3.1, GPT4 and Mistral in creating scrapers
Sep 7, 2024
•
Pierluigi Vinciguerra
1
The Lab #59: Bypassing certificate pinning with Frida and Fiddler - part 2
Rooting a virtual mobile Android Device to install Frida and discover API endpoints under the hood of apps
Aug 15, 2024
•
Pierluigi Vinciguerra
7
The Lab #58: Intercepting traffic from an App - part 1
Discover API endpoints called by an App to scrape its data
Aug 9, 2024
•
Pierluigi Vinciguerra
6
4
The Lab #57: Improving your Playwright scraper and avoid CDP detection
How to use the latest advancements to avoid CDP detection in your Playwright scrapers
Jul 27, 2024
•
Pierluigi Vinciguerra
5
3
The Lab #56: Bypassing PerimeterX 3
Testing the latest PerimeterX version by scraping Crunchbase public data
Jul 11, 2024
•
Pierluigi Vinciguerra
5
The Lab #55: Checking your browser fingerprint
Understanding your scrapers' browser fingerprint reliability with online tests.
Jul 5, 2024
•
Pierluigi Vinciguerra
9
3
The Lab #54: Scraping from Algolia APIs
Why internal APIs are always the best choice for scraping a website
Jun 21, 2024
•
Pierluigi Vinciguerra
2
2
The Lab #53: Bypassing AWS WAF
Scraping websites protected by AWS WAF using an hybrid approach
Jun 7, 2024
•
Pierluigi Vinciguerra
5
2
The Lab #52: Scraping with LLMs and ScrapeGraphAi - part 1
Are LLMs the Holy Graal for web scraping?
May 30, 2024
•
Pierluigi Vinciguerra
2
The Lab #51: APIs with Bearer Token
Scraping data from API endpoints requiring Bearer Token
May 17, 2024
•
Pierluigi Vinciguerra
5
Celebrating the 50th article of The Lab series
A brief review of the first 50 episodes of The Lab series
May 10, 2024
•
Pierluigi Vinciguerra
4
The Lab #49: Bypassing Cloudflare with open source repositories
And my two cents about these solutions
May 3, 2024
•
Pierluigi Vinciguerra
2
The Lab #48: Scraping with AWS Lambda
Using Serverless and Selenium on Lambda for gathering data
Apr 12, 2024
•
Pierluigi Vinciguerra
5
The Lab #47: Scraping real time data with Python
Using WebSocket to scrape data from Bitstamp and Sofascore
Apr 5, 2024
•
Pierluigi Vinciguerra
6
1
The Lab #46: Fingerprint injection in Playwright
A home-made solution to bypass anti-bots by changing your browser fingerprint.
Mar 29, 2024
•
Pierluigi Vinciguerra
5
THE LAB #45: Bypassing Geo-fencing While Scraping
How to scrape websites that are banned in your country
Mar 22, 2024
•
Pierluigi Vinciguerra
2
The Lab #44: Scraping the dark web
Scraping the dark web with Playwright and Brave
Mar 7, 2024
•
Pierluigi Vinciguerra
3