Repository navigation
Don't treat an embedded reCAPTCHA widget as a failed scrape - #66
Merged
Merged
Conversation
The `g-recaptcha` substring was added to `_validate_scrape` as a challenge-page marker, but the widget is ordinary page furniture on plenty of legitimate pages: every PubMed abstract page embeds one in its "Email" form. As a result `_validate_scrape` rejected perfectly good scrapes, which broke `test_journal_scraping` — both articles the cassette covers were discarded as invalid, so the scraper walked past `limit` onto a third PMID that the cassette has no response for and VCR raised CannotOverwriteExistingCassetteException. Split the patterns into ones that unambiguously mark a blocked page and the reCAPTCHA markers, and only count the latter when the page has essentially no visible text — which is what distinguishes a challenge interstitial (~165 characters) from a real article page (~10k+). The existing challenge-page fixture still trips all three of the other reCAPTCHA patterns, so it stays flagged. Add regression tests covering both directions, with a captured PubMed abstract page as the false-positive fixture. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MXqD94BZLTNQfXYM5N6Nxf
check_for_substitute_url could only build a PLoS full-text URL when it was handed one already: it read the DOI out of an `article?id=<doi>` URL. But E-utilities reports "No LinkOut links available" for plenty of PLoS articles, and `prlinks&retmode=ref` then redirects to the PubMed abstract page. The regex found no DOI there, raised, and the bare `except` returned the URL untouched -- so ACE saved the PubMed abstract page as the article. Scraping "succeeded" and ingest then found no tables, because an abstract page has none. Fall back to the DOI in the page's citation_doi meta tag, which the PubMed abstract page carries, and build the full-text URL from that. The two articles the test cassette covers now come back as PLoS NLM XML with 3 and 18 <table-wrap> elements, where before they were abstract pages. Generalize the journal check to PLoS's other titles while here, and address the current `article/file?id=<doi>` endpoint directly instead of the legacy `asset?id=<doi>.XML` one that only redirects to it. test_journal_scraping asserted nothing but a file count, so it could not see any of this; it now checks each saved document is PLoS XML. It also left its output directory behind on failure, and with index_pmids=True those leftovers push a re-run past the PMIDs the cassette covers, so every subsequent run failed for a different reason than the first. Clear the directory up front and remove it in a finally block. Cassette interactions for the PLoS fetches were recorded live. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MXqD94BZLTNQfXYM5N6Nxf
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The
g-recaptchasubstring was added to_validate_scrapeas a challenge-page marker, but the widget is ordinary page furniture on plenty of legitimate pages: every PubMed abstract page embeds one in its "Email" form. As a result_validate_scraperejected perfectly good scrapes, which broketest_journal_scraping— both articles the cassette covers were discarded as invalid, so the scraper walked pastlimitonto a third PMID that the cassette has no response for and VCR raised CannotOverwriteExistingCassetteException.Split the patterns into ones that unambiguously mark a blocked page and the reCAPTCHA markers, and only count the latter when the page has essentially no visible text — which is what distinguishes a challenge interstitial (~165 characters) from a real article page (~10k+). The existing challenge-page fixture still trips all three of the other reCAPTCHA patterns, so it stays flagged.
Add regression tests covering both directions, with a captured PubMed abstract page as the false-positive fixture.
Claude-Session: https://claude.ai/code/session_01MXqD94BZLTNQfXYM5N6Nxf