Python Async Scraper with Rotating Proxies & Rate Limit | DevPrompt Lab

Thu thập dữ liệu web quy mô lớn bằng httpx / Playwright async, xử lý Cloudflare, rotating proxies và exponential backoff.

Bạn là Web Automation & Data Engineering Specialist.

Hãy thiết kế module cào dữ liệu web an toàn bằng Python (httpx async hoặc Playwright):

Mục tiêu thu thập:

{{SCRAPE_TARGET_DESCRIPTION}}

Yêu cầu:

1. **Asynchronous Concurrency**: Dùng asyncio.Semaphore để giới hạn tối đa N requests đồng thời (tránh bị ban IP).

2. **Anti-Detect Headers**: Giả lập trình duyệt thực tế (User-Agent ngẫu nhiên hợp lệ, Sec-Ch-Ua, Accept-Language).

3. **Retry & Backoff**: Tự động thử lại khi gặp status 429 (Too Many Requests) hoặc 503 với jitter ngẫu nhiên.

4. **Data Pipeline**: Lưu dữ liệu trích xuất vào Pydantic model để xác thực trước khi lưu vào database.