Skip to content

Don't treat an embedded reCAPTCHA widget as a failed scrape - #66

Merged
jdkent merged 2 commits into
neurosynth:masterfrom
jdkent:claude/festive-cerf-m4t241
Sep 23, 2026
Merged

jdkent merged 2 commits into
neurosynth:masterfrom
jdkent:claude/festive-cerf-m4t241

Conversation

@jdkent

@jdkent jdkent commented Sep 23, 2026

Copy link
Copy Markdown
Collaborator

The g-recaptcha substring was added to _validate_scrape as a challenge-page marker, but the widget is ordinary page furniture on plenty of legitimate pages: every PubMed abstract page embeds one in its "Email" form. As a result _validate_scrape rejected perfectly good scrapes, which broke test_journal_scraping — both articles the cassette covers were discarded as invalid, so the scraper walked past limit onto a third PMID that the cassette has no response for and VCR raised CannotOverwriteExistingCassetteException.

Split the patterns into ones that unambiguously mark a blocked page and the reCAPTCHA markers, and only count the latter when the page has essentially no visible text — which is what distinguishes a challenge interstitial (~165 characters) from a real article page (~10k+). The existing challenge-page fixture still trips all three of the other reCAPTCHA patterns, so it stays flagged.

Add regression tests covering both directions, with a captured PubMed abstract page as the false-positive fixture.

Claude-Session: https://claude.ai/code/session_01MXqD94BZLTNQfXYM5N6Nxf

The `g-recaptcha` substring was added to `_validate_scrape` as a
challenge-page marker, but the widget is ordinary page furniture on
plenty of legitimate pages: every PubMed abstract page embeds one in
its "Email" form. As a result `_validate_scrape` rejected perfectly
good scrapes, which broke `test_journal_scraping` — both articles the
cassette covers were discarded as invalid, so the scraper walked past
`limit` onto a third PMID that the cassette has no response for and
VCR raised CannotOverwriteExistingCassetteException.

Split the patterns into ones that unambiguously mark a blocked page
and the reCAPTCHA markers, and only count the latter when the page has
essentially no visible text — which is what distinguishes a challenge
interstitial (~165 characters) from a real article page (~10k+). The
existing challenge-page fixture still trips all three of the other
reCAPTCHA patterns, so it stays flagged.

Add regression tests covering both directions, with a captured PubMed
abstract page as the false-positive fixture.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MXqD94BZLTNQfXYM5N6Nxf
check_for_substitute_url could only build a PLoS full-text URL when it
was handed one already: it read the DOI out of an
`article?id=<doi>` URL. But E-utilities reports "No LinkOut links
available" for plenty of PLoS articles, and `prlinks&retmode=ref` then
redirects to the PubMed abstract page. The regex found no DOI there,
raised, and the bare `except` returned the URL untouched -- so ACE
saved the PubMed abstract page as the article. Scraping "succeeded"
and ingest then found no tables, because an abstract page has none.

Fall back to the DOI in the page's citation_doi meta tag, which the
PubMed abstract page carries, and build the full-text URL from that.
The two articles the test cassette covers now come back as PLoS NLM
XML with 3 and 18 <table-wrap> elements, where before they were
abstract pages. Generalize the journal check to PLoS's other titles
while here, and address the current `article/file?id=<doi>` endpoint
directly instead of the legacy `asset?id=<doi>.XML` one that only
redirects to it.

test_journal_scraping asserted nothing but a file count, so it could
not see any of this; it now checks each saved document is PLoS XML.
It also left its output directory behind on failure, and with
index_pmids=True those leftovers push a re-run past the PMIDs the
cassette covers, so every subsequent run failed for a different
reason than the first. Clear the directory up front and remove it in
a finally block.

Cassette interactions for the PLoS fetches were recorded live.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MXqD94BZLTNQfXYM5N6Nxf
@jdkent
jdkent merged commit d64291e into neurosynth:master Sep 23, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants