Python Async Scraper with Rotating Proxies & Rate Limit | DevPrompt Lab
Thu thập dữ liệu web quy mô lớn bằng httpx / Playwright async, xử lý Cloudflare, rotating proxies và exponential backoff.
Bạn là Web Automation & Data Engineering Specialist.
Hãy thiết kế module cào dữ liệu web an toàn bằng Python (httpx async hoặc Playwright):
Mục tiêu thu thập:
{{SCRAPE_TARGET_DESCRIPTION}}
Yêu cầu:
1. **Asynchronous Concurrency**: Dùng asyncio.Semaphore để giới hạn tối đa N requests đồng thời (tránh bị ban IP).
2. **Anti-Detect Headers**: Giả lập trình duyệt thực tế (User-Agent ngẫu nhiên hợp lệ, Sec-Ch-Ua, Accept-Language).
3. **Retry & Backoff**: Tự động thử lại khi gặp status 429 (Too Many Requests) hoặc 503 với jitter ngẫu nhiên.
4. **Data Pipeline**: Lưu dữ liệu trích xuất vào Pydantic model để xác thực trước khi lưu vào database.