A collection of Python tools for retrieving, filtering, and analyzing Substack articles, including a browser-based crawler that can access Substack's built-in AI text scanner.
- Python 3
- Google Chrome (for the browser-based crawler)
- Selenium (for browser automation)
- A Substack account (for accessing the AI scanner)
Install Selenium using:
python3 -m pip install seleniumThe other scripts use Python's standard library.
| File | Description |
|---|---|
substack_tool.py |
Retrieves articles using Substack's unofficial API, with filtering and export options |
substack_crawler.py |
Uses Selenium to collect articles and interact with Substack's AI text scanner |
substack_web.py |
Provides a local web interface for collecting and filtering articles |
compare_timing.py |
Compares performance statistics from the API collector and Selenium crawler |
Clone the repository:
git clone https://github.com/YOUR_USERNAME/YOUR_REPOSITORY.git
cd YOUR_REPOSITORYInstall the required dependency:
python3 -m pip install seleniumRetrieve articles from a Substack publication:
python3 substack_tool.py fetch noahpinion --limit 25 -o articles.jsonInclude full article text:
python3 substack_tool.py fetch noahpinion --limit 25 --full-text -o articles.jsonFilter previously collected articles:
python3 substack_tool.py filter articles.json --title "AI" -o filtered.csvThe browser crawler uses Selenium to navigate Substack and interact with the Reader's "Scan for AI text" feature.
To collect articles without scanning:
python3 substack_crawler.py noahpinion --limit 5 --no-scanTo use the AI scanner, you must be logged into Substack in the Chrome browser used by Selenium.
Recommended setup for attaching to an authenticated Chrome session:
On macOS, launch Chrome with remote debugging enabled:
"/Applications/Google Chrome.app/Contents/MacOS/Google Chrome" \
--remote-debugging-port=9222 \
--user-data-dir="$HOME/substack-chrome"Log into Substack in that Chrome window.
Then, in another terminal, run:
python3 substack_crawler.py noahpinion \
--limit 5 \
--attach --show \
-o crawl.jsonThe crawler attempts to navigate from each article to its Substack Reader view, open the article's options menu, and select "Scan for AI text."
Important: Reader navigation and scanning depend on Substack's current interface. Some publications, articles, or page layouts may not work without modifications.
Launch the web application:
python3 substack_web.pyOpen http://127.0.0.1:8000 in your browser.
The interface supports fetching, filtering, sorting, and exporting collected articles.
Run the API collector and browser crawler to generate their respective timing files, then execute:
python3 compare_timing.pyThis compares execution time, average processing time per article, and other performance metrics.
Do not commit Chrome user profiles, authentication cookies, environment files containing secrets, or debug HTML captured from authenticated sessions.
Keep your Chrome remote debugging port accessible only locally.
- The API collector relies on undocumented Substack endpoints that may change.
- The browser crawler depends on Substack's current HTML structure and UI behavior.
- AI scanner results are generated by a third-party detection system and should not be treated as definitive evidence of AI authorship.
- Some content may require a subscription or additional permissions.
- Users should respect Substack's terms of service, access restrictions, and reasonable request limits.
This is an independent project and is not affiliated with or endorsed by Substack or Pangram.