A modern, reusable, and scalable web scraper for Wine Enthusiast ratings based on their Algolia API.
This scraper retrieves wine reviews from the Wine Enthusiast website's ratings section (https://www.wineenthusiast.com/ratings/) using their Algolia-based API. It's designed to be:
- Efficient: Uses direct API calls instead of HTML scraping
- Respectful: Includes configurable delays between requests
- Flexible: Supports various filtering options
- Scalable: Can handle large datasets with pagination
- Versatile: Outputs data in JSON or CSV formats
pip install -r requirements.txt
python wine_enthusiast_scraper.pyThis will scrape all available wine reviews and save them as JSON in a data directory.
python wine_enthusiast_scraper.py --output-dir custom_dir --output-format csv --max-pages 5 --delay 2.0The scraper supports various filtering options:
python wine_enthusiast_scraper.py \
--filter-country "United States" \
--filter-wine-type "Red" \
--filter-rating-min 90 \
--filter-price-max 50 \
--filter-year 2023 \
--filter-vintage 2020| Argument | Description | Default |
|---|---|---|
--output-dir |
Directory to save output files | data |
--output-format |
Output file format (json or csv) |
json |
--max-pages |
Maximum number of pages to scrape | All pages |
--delay |
Delay between requests in seconds | 1.0 |
--filter-country |
Filter by country (e.g., "United States") | None |
--filter-wine-type |
Filter by wine type (e.g., "Red", "White") | None |
--filter-rating-min |
Filter by minimum rating | None |
--filter-rating-max |
Filter by maximum rating | None |
--filter-price-min |
Filter by minimum price | None |
--filter-price-max |
Filter by maximum price | None |
--filter-year |
Filter by publication year | None |
--filter-vintage |
Filter by wine vintage year | None |
The scraper retrieves comprehensive wine review data including:
- Basic information (name, brand, vintage)
- Ratings and reviews
- Price and production details
- Region and appellation information
- Wine characteristics (varietal, type)
- Publication dates
This scraper replaces the older scrape-winemag.py which was designed for the previous website structure that required HTML parsing. The new approach offers several advantages:
- More reliable: Direct API access instead of HTML parsing
- More efficient: Retrieves structured data directly
- More complete: Captures all available metadata
- More maintainable: Simpler code structure
MIT