Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

Substack Article Collector and AI Scanner

A collection of Python tools for retrieving, filtering, and analyzing Substack articles, including a browser-based crawler that can access Substack's built-in AI text scanner.

Requirements

  • Python 3
  • Google Chrome (for the browser-based crawler)
  • Selenium (for browser automation)
  • A Substack account (for accessing the AI scanner)

Install Selenium using:

python3 -m pip install selenium

The other scripts use Python's standard library.

Files

File Description
substack_tool.py Retrieves articles using Substack's unofficial API, with filtering and export options
substack_crawler.py Uses Selenium to collect articles and interact with Substack's AI text scanner
substack_web.py Provides a local web interface for collecting and filtering articles
compare_timing.py Compares performance statistics from the API collector and Selenium crawler

Installation

Clone the repository:

git clone https://github.com/YOUR_USERNAME/YOUR_REPOSITORY.git
cd YOUR_REPOSITORY

Install the required dependency:

python3 -m pip install selenium

Usage

1. API-based article collection

Retrieve articles from a Substack publication:

python3 substack_tool.py fetch noahpinion --limit 25 -o articles.json

Include full article text:

python3 substack_tool.py fetch noahpinion --limit 25 --full-text -o articles.json

Filter previously collected articles:

python3 substack_tool.py filter articles.json --title "AI" -o filtered.csv

2. Browser-based article collection and AI scanning

The browser crawler uses Selenium to navigate Substack and interact with the Reader's "Scan for AI text" feature.

To collect articles without scanning:

python3 substack_crawler.py noahpinion --limit 5 --no-scan

To use the AI scanner, you must be logged into Substack in the Chrome browser used by Selenium.

Recommended setup for attaching to an authenticated Chrome session:

On macOS, launch Chrome with remote debugging enabled:

"/Applications/Google Chrome.app/Contents/MacOS/Google Chrome" \
  --remote-debugging-port=9222 \
  --user-data-dir="$HOME/substack-chrome"

Log into Substack in that Chrome window.

Then, in another terminal, run:

python3 substack_crawler.py noahpinion \
  --limit 5 \
  --attach --show \
  -o crawl.json

The crawler attempts to navigate from each article to its Substack Reader view, open the article's options menu, and select "Scan for AI text."

Important: Reader navigation and scanning depend on Substack's current interface. Some publications, articles, or page layouts may not work without modifications.

3. Local web interface

Launch the web application:

python3 substack_web.py

Open http://127.0.0.1:8000 in your browser.

The interface supports fetching, filtering, sorting, and exporting collected articles.

4. Performance comparison

Run the API collector and browser crawler to generate their respective timing files, then execute:

python3 compare_timing.py

This compares execution time, average processing time per article, and other performance metrics.

Security

Do not commit Chrome user profiles, authentication cookies, environment files containing secrets, or debug HTML captured from authenticated sessions.

Keep your Chrome remote debugging port accessible only locally.

Limitations

  • The API collector relies on undocumented Substack endpoints that may change.
  • The browser crawler depends on Substack's current HTML structure and UI behavior.
  • AI scanner results are generated by a third-party detection system and should not be treated as definitive evidence of AI authorship.
  • Some content may require a subscription or additional permissions.
  • Users should respect Substack's terms of service, access restrictions, and reasonable request limits.

Disclaimer

This is an independent project and is not affiliated with or endorsed by Substack or Pangram.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages