Skip to content

Latest commit

 

History

History
729 lines (577 loc) · 23.3 KB

File metadata and controls

729 lines (577 loc) · 23.3 KB

Python API

One call

from citeprobe import verify

result = verify(
    "Vaswani et al. (2017). Attention Is All You Need. "
    "NeurIPS. arXiv:1706.03762"
)

print(result.verdict)
print(result.to_dict())

verify() accepts a string or a parsed ReferenceQuery. Set cache=False for an uncached run, or pass a custom object that follows the ResultCache protocol.

Async code

from citeprobe import verify_async

result = await verify_async("doi:10.1038/nphys1170")

Do not call the synchronous function from an active event loop. Citeprobe raises a direct error that points to verify_async().

Default adapters accept retry controls through the one-call API:

result = verify(
    "doi:10.1038/nphys1170",
    retries=4,
    retry_backoff=1.0,
)

Custom adapters can use RequestPolicy with request_source().

Parse files

from citeprobe import (
    load_references,
    parse_numbered_references,
    parse_reference,
    parse_reference_text,
)

one = parse_reference("arXiv:1706.03762")
many = load_references("references.bib")
numbered = parse_numbered_references("[1] First work\n[2] Second work")
detected = parse_reference_text("[1] First work\n[2] Second work")

Pass max_references=2000 to stop an importer when a file exceeds an application-level record limit. parse_reference_text() uses guarded numbered block detection by default. Set input_format="lines" or input_format="numbered" to choose the interpretation directly.

Structured bibliography parsing is local. PDF input uses GROBID by default:

from citeprobe import extract_pdf_references

queries = extract_pdf_references(
    "paper.pdf",
    grobid_url="http://127.0.0.1:8070",
    consolidation=0,
    timeout=120,
)

The optional text-layer backend runs without a service:

queries = extract_pdf_references(
    "paper.pdf",
    backend="local",
)

Use an extraction report when the selected method and layout decisions need to remain visible:

from citeprobe import extract_pdf_report

report = extract_pdf_report("paper.pdf", backend="auto")
print(report.requested_backend)
print(report.backend)
print(report.fallback_used)
print(report.pages_read)
print(report.bibliography_page)
print(report.layout)
print(report.warnings)
print(report.consolidation)
print(report.include_raw)
print(len(report.references))

Install citeprobe[pdf] for that backend. It supports born-digital PDFs but does not perform OCR. Select OCR explicitly for scanned papers:

from citeprobe import extract_pdf_report, ocr_capabilities

print(ocr_capabilities())
report = extract_pdf_report(
    "scan.pdf",
    backend="ocr",
    ocr_languages="eng+deu",
    ocr_dpi=300,
    ocr_page_segmentation_mode=1,
)
print(report.ocr_engine_version)
print(report.ocr_mean_confidence)

OCR renders pages with Poppler and recognizes them with Tesseract. Both tools must be on PATH. Language packs named by ocr_languages must also be installed. DPI is bounded from 150 to 600, and Tesseract page segmentation mode is bounded from 0 to 13. Temporary page images are deleted after the run.

backend="auto" tries local segmentation first and can upload to GROBID after a segmentation-confidence failure. It does not silently run OCR. Other local errors do not trigger an upload. GROBID consolidation is accepted only with backend="grobid".

Async applications can keep extraction off the event loop:

from citeprobe import extract_pdf_references_async

queries = await extract_pdf_references_async(
    "paper.pdf",
    grobid_url="http://127.0.0.1:8070",
    consolidation=0,
)

extract_pdf_report_async() returns the same provenance fields while keeping local extraction off the event loop.

GrobidClient exposes PDF and response size limits for an embedded application. Consolidation accepts 0, 1, or 2; mode 0 avoids consolidation calls made by the GROBID service. Modes 1 and 2 set ReferenceQuery.input_enrichment, which prevents a later verification from returning verified. Verification sends the extracted metadata to the selected Citeprobe providers.

Remote GROBID URLs require HTTPS. An application can set allow_insecure_remote=True only for a trusted plain-HTTP deployment.

Select sources

from citeprobe import CrossrefSource, DataCiteSource, OpenAlexSource, verify

result = verify(
    "doi:10.1038/nphys1170",
    sources=(CrossrefSource(), DataCiteSource(), OpenAlexSource()),
)

Source tuples must be nonempty and source names must be unique.

The registry exposes all built-in adapters:

from citeprobe import SOURCE_NAMES, build_sources, source_capabilities

print(SOURCE_NAMES)
sources = build_sources("crossref,openalex,europe_pmc")
for item in source_capabilities():
    print(item.name, item.direct_identifiers, item.author_coverage)

Eight scholarly sources are enabled by default. Specialist legal, government, standards, security-preprint, book, and local-index adapters are opt-in because their applicability or credentials differ.

Select checks

from citeprobe import verify

result = verify(
    "doi:10.1038/nphys1170",
    checks=("title", "authors", "year"),
)

Pass "all" or omit checks to run every registered check. The CHECK_NAMES tuple and CHECK_DESCRIPTIONS mapping expose the current registry. Each result stores the checks that ran in result.checks.

CandidateRecord.publication_status records accepted, rejected, desk-rejected, withdrawn, submitted, or unknown state. OpenReview supplies this field when a native API2 note matches the public venue group's status IDs or a native API1 note has a public direct decision reply. Other adapters and imported catalog copies leave it unknown, so the publication-status check skips them. A failed OpenReview status request leaves the field unknown, marks the source run as an error, and degrades this check. Conflicting known statuses also become unknown.

Verifier(checks=...) disables OpenReview status requests when it constructs the default sources and the check is not selected. When supplying a custom source tuple, construct OpenReviewSource(publication_status=False) to get the same behavior.

VerificationResult.suggestions contains proposed additions and replacements. Each FieldSuggestion records the cited and proposed value, its evidence basis, decision, reason, provider names, candidate IDs, shared identity, and independent evidence groups. Suggestions are read-only. Citeprobe abstains when the matched records disagree or do not meet the evidence rule. Pass suggest_missing=True to verify(), verify_async(), verify_query(), or Verifier() to include optional additions.

Corrected BibTeX

from citeprobe import render_corrected_bibtex

exported = render_corrected_bibtex((result,))
print(exported.entry_count)
print(exported.applied_count)
print(exported.review_count)
print(exported.text)

render_corrected_bibtex() applies only proposals marked suggest. Review proposals remain one-line JSON comments and never change fields. Pass include_review_comments=False to omit those comments. The function returns a BibtexExport value and does not write a file.

BibTeX imports carry a typed BibtexOrigin in each ReferenceQuery. It keeps the parsed entry type, key, fields, and person roles so generated output does not drop editors or unknown fields. A file-level preamble is stored on the first parsed query and emitted once for the batch. Source layout is normalized, string macros are expanded, and source comments are not retained. Queries from other formats become canonical @misc entries.

Duplicate checks need a batch. Verifier.verify_many() adds the findings to each affected result:

from citeprobe import Verifier, load_references

queries = load_references("references.bib")
results = Verifier().verify_many(queries)

The detector can also run without source requests:

from citeprobe import find_duplicate_references

matches = find_duplicate_references(
    queries,
    near_title=0.90,
)

This local path needs no API, model, or special hardware. It uses normalized identifiers, titles, first-author families, and years. Near-title comparisons are blocked by first-author family. Large author blocks use necessary token, length, character-overlap, and Indel bounds before the more expensive title score. The bounds start above 64 entries by default. max_block_size changes that point without changing the score rule. Near-title detection remains exhaustive, so its worst-case pair enumeration is quadratic.

URL checks block private and local targets by default. A trusted local target can be enabled for one call:

result = verify(
    "http://127.0.0.1:8000/paper",
    checks="url",
    allow_private_urls=True,
)

For a network with a fixed DNS64 prefix, pass a configured checker:

from citeprobe import UrlChecker, Verifier

verifier = Verifier(
    url_checker=UrlChecker(
        nat64_prefixes=("2001:db8:64::/96",),
    )
)

Use a custom policy

from citeprobe import VerificationPolicy, verify

policy = VerificationPolicy(
    name="default",
    title_match=0.9,
    title_contradiction=0.55,
    year_tolerance=0,
    minimum_sources=2,
    missing_author_sources=2,
    allow_single_authority=False,
)

result = verify("A pasted reference", policy=policy)

The policy name field currently uses the built-in policy-name type. Custom thresholds can still be supplied through a VerificationPolicy instance. minimum_sources and missing_author_sources count independent evidence groups, not adapter names. A source adapter can declare SourceCapabilities.default_evidence_group when it mirrors or proxies another catalog. Candidate output retains both the provider name and assigned group.

Define source trust

A custom adapter's SourceRun.source and every CandidateRecord.source must equal SourceCapabilities.name. Citeprobe rejects the whole source result when those identities differ.

Set authoritative_identifiers only for identifier types whose primary registration metadata the adapter returns. Each declared authority must also appear in direct_identifiers. A mirror can share an evidence group without receiving the original provider's authority.

default_evidence_group applies to every candidate unless that candidate sets evidence_group. Candidate overrides are for adapters that return records from more than one upstream catalog. The default group and authority declarations are included in the cache identity. Bump cache_namespace when the logic that assigns candidate-level groups changes.

Run a parsed batch

from citeprobe import Verifier, load_references

queries = load_references("references.bib")
results = Verifier().verify_many(queries, max_concurrency=4)

The batch preserves input order and reuses one HTTP client. Each reference still receives its own result, candidates, findings, and source runs.

Load a reference archive

from citeprobe import ArchiveLimits, Verifier, load_reference_archive

archive = load_reference_archive(
    "reference-library.zip",
    member_types=("bib", "ris", "pdf"),
    max_references_per_member=500,
    max_total_references=1000,
    pdf_backend="local",
    limits=ArchiveLimits(
        max_archive_bytes=50 * 1024 * 1024,
        max_members=2000,
        max_member_bytes=25 * 1024 * 1024,
        max_total_bytes=200 * 1024 * 1024,
        max_metadata_bytes=8 * 1024 * 1024,
        max_path_bytes=4096,
        max_compression_ratio=200,
    ),
)
results = Verifier().verify_many(archive.queries, max_concurrency=4)

ReferenceArchiveReport.members records each parsed, skipped, or failed member. references pairs every query with its member path. The report also returns the archive format, compressed byte count, expanded byte count, SHA-256, and selected member types.

By default, a selected member parse error is recorded and later members are still parsed. Set fail_fast=True to stop on the first one. Archive-wide limits always stop loading. Temporary extraction files are removed before the function returns or raises.

Archived URL fallback is opt-in:

from citeprobe import verify

result = verify(
    "A. Author. Example. https://example.com/missing",
    checks="url",
    wayback=True,
)

A valid snapshot affects only the URL check. Its URL and capture details are stored in the URL source run and do not count as metadata-source agreement.

Read evidence

for finding in result.findings:
    print(finding.code, finding.severity, finding.field, finding.source)

for check in result.check_results:
    print(check.check, check.status, check.evidence_sources)

for author in result.author_evidence:
    print(author.cited, author.status, author.matched, author.reason)

for run in result.source_runs:
    print(run.source, run.status, run.elapsed_ms, run.cached)

Public result models are frozen dataclasses. Tuples are used for ordered collections so a completed result cannot change under the caller.

Offline evidence

The one-call API can use existing cache entries without network requests:

from citeprobe import DiskCache, verify

result = verify(
    "doi:10.1038/nphys1170",
    cache=DiskCache(".citeprobe-cache"),
    offline=True,
)

A missing cache entry remains unavailable evidence. For exact replay, write a versioned snapshot after a run:

from citeprobe import (
    __version__,
    read_evidence_snapshot,
    write_evidence_snapshot,
)

digest = write_evidence_snapshot(
    "evidence.json",
    tuple(results),
    citeprobe_version=__version__,
    input_metadata={"kind": "bibtex", "name": "references.bib"},
)
snapshot = read_evidence_snapshot("evidence.json")
assert snapshot.artifact_sha256 == digest
replayed = snapshot.replay()

The reader enforces a byte limit, rejects unknown or duplicate fields, rebuilds the typed result graph, and validates the embedded SHA-256. Replay returns those immutable results and does not create a source adapter.

Saved review decisions

from citeprobe import ReviewStore

store = ReviewStore("review.sqlite3")
stored = store.save_run(
    results,
    label="journal resubmission",
    options={"policy": "strict"},
)
update = store.update_decision(
    stored.run_id,
    reference_index=0,
    finding_code="AUT201",
    finding_index=0,
    disposition="false_positive",
    note="Verified source spelling",
    propagation="save",
    expected_run_instance_id=stored.run_instance_id,
    expected_decision_revision="absent",
    expected_annotation_revision="absent",
)
print(update.annotation_changed, update.annotation_id)
print(store.get_run(stored.run_id))
print(store.decisions(stored.run_id))
print(store.matched_decisions(stored.run_id))
print(store.get_review_snapshot(stored.run_id))

Saved result sets are immutable and digest-checked when read. A reviewer decision is stored separately, can be changed, and never rewrites source evidence. list_runs(), delete_run(), and prune() provide bounded history management.

StoredRun.run_instance_id is an opaque mutation token. Pass it as expected_run_instance_id to set_decision(), update_decision(), and delete_run(). If the run was deleted and its public ID was reused, the write raises ReviewConflictError before reading or changing the replacement.

propagation accepts none, save, or remove. Saving a cross-run annotation requires an exact cited identifier corroborated by a resolved candidate. Matching also includes the finding code, field, source, stable subject, and result schema version. Only REF findings plus AUT201, AUT204, and AUT205 can propagate. Operational source errors cannot. matched_decisions() returns direct decisions first and then eligible annotations for findings without a direct decision.

annotation_conflicts() returns disagreeing exact-ID rule sets without inventing an effective disposition. review_state() returns the matched and conflicting tuples from one database snapshot. Each AnnotationConflict contains the finding position, rule IDs, dispositions, and annotation revision needed to replace or remove the set.

get_review_snapshot() returns the immutable run, direct decisions, effective decisions, and conflicts from one explicit SQLite read transaction. It returns None when the run does not exist. The local UI detail API and runs show use this aggregate read, so one response cannot mix states from concurrent writes.

update_decision() returns a ReviewDecisionUpdate. Its annotation_changed field distinguishes a real remove from a no-op, and annotation_id identifies a newly saved rule. set_decision() remains the short form when only the stored ReviewDecision is needed.

Use finding_index or finding_subject to select repeated findings. decision_revision and annotation_revision are opaque optimistic-lock tokens returned by decision reads. Pass them back as expected_decision_revision and expected_annotation_revision, or pass absent to require no existing record. ReviewConflictError means another writer changed the selected state. The transaction is rolled back. Annotation revisions may guard a run-only decision, which prevents stale inherited state from becoming a new direct decision.

Rules with several identifiers are joined only when one resolved candidate corroborates those IDs together. Conflicting same-kind IDs in the complete citation cause abstention, including when only one value resolves. Conflicting matched dispositions are returned as AnnotationConflict.

Annotations retain their immutable rule ID, source result digest, source reference position, and source run-instance ID. They survive deletion or pruning of their source run and identify run-ID reuse even when the replacement results are identical. Migrated legacy rules without an original instance remain unknown. They never replace the raw VerificationResult, change a verdict, or alter report and exit behavior.

Review database schema 5 adds run-instance provenance. Migration is atomic, rejects orphan legacy decisions, and does not infer missing provenance from a run that merely has the same public ID.

Local metadata index

from citeprobe import (
    LocalIndexSource,
    build_local_index,
    create_local_index,
    inspect_local_index,
)

create_local_index("records.db", provenance="2026-07 evidence snapshots")
candidates = tuple(
    candidate
    for result in results
    for candidate in result.candidates
)
build_local_index("records.db", candidates)
print(inspect_local_index("records.db"))

local_result = verify(
    "doi:10.1038/nphys1170",
    sources=(LocalIndexSource("records.db"),),
)

The index uses normalized identifiers for exact lookup and FTS title search as a fallback. Its provenance string is attached to returned candidates. Creating and building are separate from the read-only adapter so verification cannot alter the index.

Manuscript structure

from citeprobe import inspect_manuscript, select_manuscript_checks

report = inspect_manuscript("paper.tex", project_root="latex-project")
focused = select_manuscript_checks(
    report,
    ("undefined", "uncited", "duplicates"),
)
for finding in focused.findings:
    print(finding.code, finding.anchor)

inspect_manuscript() accepts LaTeX, DOCX, ODT, RTF, and Markdown. A LaTeX directory is also accepted and gets a deterministic main-file selection. Literal local includes and bibliography resources are followed within the project root. Dynamic paths produce an abstention instead of execution. Archive-backed office inputs have entry, expansion, and compression-ratio bounds.

The five finding families are undefined, uncited, duplicates, dynamic, and missing. The report retains source anchors, citation occurrences, bibliography keys, visited files, and an explicit complete, partial, or abstained status.

Work families

from citeprobe import WorkFamilyBuilder

builder = WorkFamilyBuilder()
for candidate in result.candidates:
    builder.add_candidate(candidate)
family = builder.build()
print(family.to_dict())

Strong identifiers merge records that describe the same work. Provider relations can add version, preprint, correction, retraction, dataset, and software edges. Title similarity can produce a relation suggestion, but it does not merge identities on its own. Unsupported or incomplete provider relations remain explicit abstentions.

LaTeX revision analysis

Compare bibliographies by key or identifier-backed identity:

from pathlib import Path
from citeprobe import semantic_bibtex_diff, semantic_citation_diff

bib_diff = semantic_bibtex_diff(
    Path("old.bib").read_text(encoding="utf-8"),
    Path("new.bib").read_text(encoding="utf-8"),
    match_mode="identifier",
)
cite_diff = semantic_citation_diff(
    "old-project",
    "new-project",
    old_main="paper.tex",
    new_main="paper.tex",
)

The bibliography result separates added, removed, modified, renamed, and unchanged entries. Rename detection needs a unique shared DOI, PMID, PMCID, arXiv ID, or ISBN. Citation comparison follows bounded literal includes and reports added, removed, and retargeted sites.

Visual diff planning keeps execution policy explicit:

from citeprobe import execute_visual_diff, plan_visual_diff

plan = plan_visual_diff(
    "old-project",
    "new-project",
    "diff-work",
    mode="source-only",
    old_main="paper.tex",
    new_main="paper.tex",
    latexdiff_profile="safe",
    old_number_mode="auto",
    addition_color="teal",
    deletion_color="magenta",
)
result = execute_visual_diff(plan)
for artifact in result.artifacts:
    print(artifact.path, artifact.sha256)

source-only runs latexdiff without TeX compilation. Native compilation requires trust_project=True. Docker mode uses a no-network container with resource limits and requires a digest-pinned image by default. Select tex_engine="pdflatex", "xelatex", or "lualatex". safe_extract_zip() provides bounded archive extraction for uploaded projects. Plans bind their execution fields and commands to an integrity digest. Compilation can retry without strike lines, and strikethrough_fallback=False disables that retry.

Benchmark manifests

from citeprobe import (
    load_benchmark_manifest,
    manifest_template,
    run_benchmark_manifest,
)

print(manifest_template())
manifest = load_benchmark_manifest("corpus.json")
outcome = run_benchmark_manifest(manifest)
print(outcome["passed"])

Every artifact has an HTTPS URL, expected SHA-256, license declaration, backend, run count, and optional ground-truth count. The runner verifies the download before extraction and removes its temporary directory after each case. Default network requests reject local and private targets, resolve each redirect again, and pin the selected public address.

Render reports

from citeprobe import render_results

markdown = render_results(results, "markdown")
sarif = render_results(results, "sarif")

Supported values are available in REPORT_FORMATS. The renderer returns text and does not write a file.

Embed the local app

The UI is an optional FastAPI application:

from citeprobe.ui import create_app

app = create_app(
    grobid_url="http://127.0.0.1:8070",
    grobid_timeout=120,
)

Install citeprobe[ui] first. The packaged command binds the app to 127.0.0.1; embedding code is responsible for its own server binding. The GROBID URL is a server setting. A PDF upload can choose consolidation mode, but it cannot supply a replacement URL.