from citeprobe import verify
result = verify(
"Vaswani et al. (2017). Attention Is All You Need. "
"NeurIPS. arXiv:1706.03762"
)
print(result.verdict)
print(result.to_dict())verify() accepts a string or a parsed ReferenceQuery. Set cache=False for
an uncached run, or pass a custom object that follows the ResultCache
protocol.
from citeprobe import verify_async
result = await verify_async("doi:10.1038/nphys1170")Do not call the synchronous function from an active event loop. Citeprobe
raises a direct error that points to verify_async().
Default adapters accept retry controls through the one-call API:
result = verify(
"doi:10.1038/nphys1170",
retries=4,
retry_backoff=1.0,
)Custom adapters can use RequestPolicy with request_source().
from citeprobe import (
load_references,
parse_numbered_references,
parse_reference,
parse_reference_text,
)
one = parse_reference("arXiv:1706.03762")
many = load_references("references.bib")
numbered = parse_numbered_references("[1] First work\n[2] Second work")
detected = parse_reference_text("[1] First work\n[2] Second work")Pass max_references=2000 to stop an importer when a file exceeds an
application-level record limit. parse_reference_text() uses guarded numbered
block detection by default. Set input_format="lines" or
input_format="numbered" to choose the interpretation directly.
Structured bibliography parsing is local. PDF input uses GROBID by default:
from citeprobe import extract_pdf_references
queries = extract_pdf_references(
"paper.pdf",
grobid_url="http://127.0.0.1:8070",
consolidation=0,
timeout=120,
)The optional text-layer backend runs without a service:
queries = extract_pdf_references(
"paper.pdf",
backend="local",
)Use an extraction report when the selected method and layout decisions need to remain visible:
from citeprobe import extract_pdf_report
report = extract_pdf_report("paper.pdf", backend="auto")
print(report.requested_backend)
print(report.backend)
print(report.fallback_used)
print(report.pages_read)
print(report.bibliography_page)
print(report.layout)
print(report.warnings)
print(report.consolidation)
print(report.include_raw)
print(len(report.references))Install citeprobe[pdf] for that backend. It supports born-digital PDFs but
does not perform OCR. Select OCR explicitly for scanned papers:
from citeprobe import extract_pdf_report, ocr_capabilities
print(ocr_capabilities())
report = extract_pdf_report(
"scan.pdf",
backend="ocr",
ocr_languages="eng+deu",
ocr_dpi=300,
ocr_page_segmentation_mode=1,
)
print(report.ocr_engine_version)
print(report.ocr_mean_confidence)OCR renders pages with Poppler and recognizes them with Tesseract. Both tools
must be on PATH. Language packs named by ocr_languages must also be
installed. DPI is bounded from 150 to 600, and Tesseract page segmentation
mode is bounded from 0 to 13. Temporary page images are deleted after the
run.
backend="auto" tries local segmentation first and can upload to GROBID after
a segmentation-confidence failure. It does not silently run OCR. Other local
errors do not trigger an upload. GROBID consolidation is accepted only with
backend="grobid".
Async applications can keep extraction off the event loop:
from citeprobe import extract_pdf_references_async
queries = await extract_pdf_references_async(
"paper.pdf",
grobid_url="http://127.0.0.1:8070",
consolidation=0,
)extract_pdf_report_async() returns the same provenance fields while keeping
local extraction off the event loop.
GrobidClient exposes PDF and response size limits for an embedded
application. Consolidation accepts 0, 1, or 2; mode 0 avoids
consolidation calls made by the GROBID service. Modes 1 and 2 set
ReferenceQuery.input_enrichment, which prevents a later verification from
returning verified. Verification sends the extracted metadata to the
selected Citeprobe providers.
Remote GROBID URLs require HTTPS. An application can set
allow_insecure_remote=True only for a trusted plain-HTTP deployment.
from citeprobe import CrossrefSource, DataCiteSource, OpenAlexSource, verify
result = verify(
"doi:10.1038/nphys1170",
sources=(CrossrefSource(), DataCiteSource(), OpenAlexSource()),
)Source tuples must be nonempty and source names must be unique.
The registry exposes all built-in adapters:
from citeprobe import SOURCE_NAMES, build_sources, source_capabilities
print(SOURCE_NAMES)
sources = build_sources("crossref,openalex,europe_pmc")
for item in source_capabilities():
print(item.name, item.direct_identifiers, item.author_coverage)Eight scholarly sources are enabled by default. Specialist legal, government, standards, security-preprint, book, and local-index adapters are opt-in because their applicability or credentials differ.
from citeprobe import verify
result = verify(
"doi:10.1038/nphys1170",
checks=("title", "authors", "year"),
)Pass "all" or omit checks to run every registered check. The
CHECK_NAMES tuple and CHECK_DESCRIPTIONS mapping expose the current
registry. Each result stores the checks that ran in result.checks.
CandidateRecord.publication_status records accepted, rejected, desk-rejected,
withdrawn, submitted, or unknown state. OpenReview supplies this field when a
native API2 note matches the public venue group's status IDs or a native API1
note has a public direct decision reply. Other adapters and imported catalog
copies leave it unknown, so the publication-status check skips them. A failed
OpenReview status request leaves the field unknown, marks the source run as an
error, and degrades this check. Conflicting known statuses also become unknown.
Verifier(checks=...) disables OpenReview status requests when it constructs
the default sources and the check is not selected. When supplying a custom
source tuple, construct OpenReviewSource(publication_status=False) to get the
same behavior.
VerificationResult.suggestions contains proposed additions and replacements.
Each FieldSuggestion records the cited and proposed value, its evidence
basis, decision, reason, provider names, candidate IDs, shared identity, and
independent evidence groups. Suggestions are read-only. Citeprobe abstains
when the matched records disagree or do not meet the evidence rule. Pass
suggest_missing=True to verify(), verify_async(), verify_query(), or
Verifier() to include optional additions.
from citeprobe import render_corrected_bibtex
exported = render_corrected_bibtex((result,))
print(exported.entry_count)
print(exported.applied_count)
print(exported.review_count)
print(exported.text)render_corrected_bibtex() applies only proposals marked suggest. Review
proposals remain one-line JSON comments and never change fields. Pass
include_review_comments=False to omit those comments. The function returns a
BibtexExport value and does not write a file.
BibTeX imports carry a typed BibtexOrigin in each ReferenceQuery. It keeps
the parsed entry type, key, fields, and person roles so generated output does
not drop editors or unknown fields. A file-level preamble is stored on the
first parsed query and emitted once for the batch. Source layout is normalized,
string macros are expanded, and source comments are not retained. Queries from
other formats become canonical @misc entries.
Duplicate checks need a batch. Verifier.verify_many() adds the findings to
each affected result:
from citeprobe import Verifier, load_references
queries = load_references("references.bib")
results = Verifier().verify_many(queries)The detector can also run without source requests:
from citeprobe import find_duplicate_references
matches = find_duplicate_references(
queries,
near_title=0.90,
)This local path needs no API, model, or special hardware. It uses normalized
identifiers, titles, first-author families, and years. Near-title comparisons
are blocked by first-author family. Large author blocks use necessary token,
length, character-overlap, and Indel bounds before the more expensive title
score. The bounds start above 64 entries by default. max_block_size changes
that point without changing the score rule. Near-title detection remains
exhaustive, so its worst-case pair enumeration is quadratic.
URL checks block private and local targets by default. A trusted local target can be enabled for one call:
result = verify(
"http://127.0.0.1:8000/paper",
checks="url",
allow_private_urls=True,
)For a network with a fixed DNS64 prefix, pass a configured checker:
from citeprobe import UrlChecker, Verifier
verifier = Verifier(
url_checker=UrlChecker(
nat64_prefixes=("2001:db8:64::/96",),
)
)from citeprobe import VerificationPolicy, verify
policy = VerificationPolicy(
name="default",
title_match=0.9,
title_contradiction=0.55,
year_tolerance=0,
minimum_sources=2,
missing_author_sources=2,
allow_single_authority=False,
)
result = verify("A pasted reference", policy=policy)The policy name field currently uses the built-in policy-name type. Custom
thresholds can still be supplied through a VerificationPolicy instance.
minimum_sources and missing_author_sources count independent evidence
groups, not adapter names. A source adapter can declare
SourceCapabilities.default_evidence_group when it mirrors or proxies another
catalog. Candidate output retains both the provider name and assigned group.
A custom adapter's SourceRun.source and every CandidateRecord.source must
equal SourceCapabilities.name. Citeprobe rejects the whole source result
when those identities differ.
Set authoritative_identifiers only for identifier types whose primary
registration metadata the adapter returns. Each declared authority must also
appear in direct_identifiers. A mirror can share an evidence group without
receiving the original provider's authority.
default_evidence_group applies to every candidate unless that candidate
sets evidence_group. Candidate overrides are for adapters that return
records from more than one upstream catalog. The default group and authority
declarations are included in the cache identity. Bump cache_namespace when
the logic that assigns candidate-level groups changes.
from citeprobe import Verifier, load_references
queries = load_references("references.bib")
results = Verifier().verify_many(queries, max_concurrency=4)The batch preserves input order and reuses one HTTP client. Each reference still receives its own result, candidates, findings, and source runs.
from citeprobe import ArchiveLimits, Verifier, load_reference_archive
archive = load_reference_archive(
"reference-library.zip",
member_types=("bib", "ris", "pdf"),
max_references_per_member=500,
max_total_references=1000,
pdf_backend="local",
limits=ArchiveLimits(
max_archive_bytes=50 * 1024 * 1024,
max_members=2000,
max_member_bytes=25 * 1024 * 1024,
max_total_bytes=200 * 1024 * 1024,
max_metadata_bytes=8 * 1024 * 1024,
max_path_bytes=4096,
max_compression_ratio=200,
),
)
results = Verifier().verify_many(archive.queries, max_concurrency=4)ReferenceArchiveReport.members records each parsed, skipped, or failed
member. references pairs every query with its member path. The report also
returns the archive format, compressed byte count, expanded byte count,
SHA-256, and selected member types.
By default, a selected member parse error is recorded and later members are
still parsed. Set fail_fast=True to stop on the first one. Archive-wide
limits always stop loading. Temporary extraction files are removed before the
function returns or raises.
Archived URL fallback is opt-in:
from citeprobe import verify
result = verify(
"A. Author. Example. https://example.com/missing",
checks="url",
wayback=True,
)A valid snapshot affects only the URL check. Its URL and capture details are stored in the URL source run and do not count as metadata-source agreement.
for finding in result.findings:
print(finding.code, finding.severity, finding.field, finding.source)
for check in result.check_results:
print(check.check, check.status, check.evidence_sources)
for author in result.author_evidence:
print(author.cited, author.status, author.matched, author.reason)
for run in result.source_runs:
print(run.source, run.status, run.elapsed_ms, run.cached)Public result models are frozen dataclasses. Tuples are used for ordered collections so a completed result cannot change under the caller.
The one-call API can use existing cache entries without network requests:
from citeprobe import DiskCache, verify
result = verify(
"doi:10.1038/nphys1170",
cache=DiskCache(".citeprobe-cache"),
offline=True,
)A missing cache entry remains unavailable evidence. For exact replay, write a versioned snapshot after a run:
from citeprobe import (
__version__,
read_evidence_snapshot,
write_evidence_snapshot,
)
digest = write_evidence_snapshot(
"evidence.json",
tuple(results),
citeprobe_version=__version__,
input_metadata={"kind": "bibtex", "name": "references.bib"},
)
snapshot = read_evidence_snapshot("evidence.json")
assert snapshot.artifact_sha256 == digest
replayed = snapshot.replay()The reader enforces a byte limit, rejects unknown or duplicate fields, rebuilds the typed result graph, and validates the embedded SHA-256. Replay returns those immutable results and does not create a source adapter.
from citeprobe import ReviewStore
store = ReviewStore("review.sqlite3")
stored = store.save_run(
results,
label="journal resubmission",
options={"policy": "strict"},
)
update = store.update_decision(
stored.run_id,
reference_index=0,
finding_code="AUT201",
finding_index=0,
disposition="false_positive",
note="Verified source spelling",
propagation="save",
expected_run_instance_id=stored.run_instance_id,
expected_decision_revision="absent",
expected_annotation_revision="absent",
)
print(update.annotation_changed, update.annotation_id)
print(store.get_run(stored.run_id))
print(store.decisions(stored.run_id))
print(store.matched_decisions(stored.run_id))
print(store.get_review_snapshot(stored.run_id))Saved result sets are immutable and digest-checked when read. A reviewer
decision is stored separately, can be changed, and never rewrites source
evidence. list_runs(), delete_run(), and prune() provide bounded history
management.
StoredRun.run_instance_id is an opaque mutation token. Pass it as
expected_run_instance_id to set_decision(), update_decision(), and
delete_run(). If the run was deleted and its public ID was reused, the write
raises ReviewConflictError before reading or changing the replacement.
propagation accepts none, save, or remove. Saving a cross-run
annotation requires an exact cited identifier corroborated by a resolved
candidate. Matching also includes the finding code, field, source, stable
subject, and result schema version. Only REF findings plus AUT201,
AUT204, and AUT205 can propagate. Operational source errors cannot.
matched_decisions() returns direct decisions first and then eligible
annotations for findings without a direct decision.
annotation_conflicts() returns disagreeing exact-ID rule sets without
inventing an effective disposition. review_state() returns the matched and
conflicting tuples from one database snapshot. Each AnnotationConflict
contains the finding position, rule IDs, dispositions, and annotation
revision needed to replace or remove the set.
get_review_snapshot() returns the immutable run, direct decisions, effective
decisions, and conflicts from one explicit SQLite read transaction. It returns
None when the run does not exist. The local UI detail API and runs show use
this aggregate read, so one response cannot mix states from concurrent writes.
update_decision() returns a ReviewDecisionUpdate. Its
annotation_changed field distinguishes a real remove from a no-op, and
annotation_id identifies a newly saved rule. set_decision() remains the
short form when only the stored ReviewDecision is needed.
Use finding_index or finding_subject to select repeated findings.
decision_revision and annotation_revision are opaque optimistic-lock
tokens returned by decision reads. Pass them back as
expected_decision_revision and expected_annotation_revision, or pass
absent to require no existing record. ReviewConflictError means another
writer changed the selected state. The transaction is rolled back.
Annotation revisions may guard a run-only decision, which prevents stale
inherited state from becoming a new direct decision.
Rules with several identifiers are joined only when one resolved candidate
corroborates those IDs together. Conflicting same-kind IDs in the complete
citation cause abstention, including when only one value resolves. Conflicting
matched dispositions are returned as AnnotationConflict.
Annotations retain their immutable rule ID, source result digest, source
reference position, and source run-instance ID. They survive deletion or
pruning of their source run and identify run-ID reuse even when the replacement
results are identical. Migrated legacy rules without an original instance
remain unknown. They never replace the raw VerificationResult, change a
verdict, or alter report and exit behavior.
Review database schema 5 adds run-instance provenance. Migration is atomic, rejects orphan legacy decisions, and does not infer missing provenance from a run that merely has the same public ID.
from citeprobe import (
LocalIndexSource,
build_local_index,
create_local_index,
inspect_local_index,
)
create_local_index("records.db", provenance="2026-07 evidence snapshots")
candidates = tuple(
candidate
for result in results
for candidate in result.candidates
)
build_local_index("records.db", candidates)
print(inspect_local_index("records.db"))
local_result = verify(
"doi:10.1038/nphys1170",
sources=(LocalIndexSource("records.db"),),
)The index uses normalized identifiers for exact lookup and FTS title search as a fallback. Its provenance string is attached to returned candidates. Creating and building are separate from the read-only adapter so verification cannot alter the index.
from citeprobe import inspect_manuscript, select_manuscript_checks
report = inspect_manuscript("paper.tex", project_root="latex-project")
focused = select_manuscript_checks(
report,
("undefined", "uncited", "duplicates"),
)
for finding in focused.findings:
print(finding.code, finding.anchor)inspect_manuscript() accepts LaTeX, DOCX, ODT, RTF, and Markdown. A LaTeX
directory is also accepted and gets a deterministic main-file selection.
Literal local includes and bibliography resources are followed within the
project root. Dynamic paths produce an abstention instead of execution.
Archive-backed office inputs have entry, expansion, and compression-ratio
bounds.
The five finding families are undefined, uncited, duplicates, dynamic,
and missing. The report retains source anchors, citation occurrences,
bibliography keys, visited files, and an explicit complete, partial, or
abstained status.
from citeprobe import WorkFamilyBuilder
builder = WorkFamilyBuilder()
for candidate in result.candidates:
builder.add_candidate(candidate)
family = builder.build()
print(family.to_dict())Strong identifiers merge records that describe the same work. Provider relations can add version, preprint, correction, retraction, dataset, and software edges. Title similarity can produce a relation suggestion, but it does not merge identities on its own. Unsupported or incomplete provider relations remain explicit abstentions.
Compare bibliographies by key or identifier-backed identity:
from pathlib import Path
from citeprobe import semantic_bibtex_diff, semantic_citation_diff
bib_diff = semantic_bibtex_diff(
Path("old.bib").read_text(encoding="utf-8"),
Path("new.bib").read_text(encoding="utf-8"),
match_mode="identifier",
)
cite_diff = semantic_citation_diff(
"old-project",
"new-project",
old_main="paper.tex",
new_main="paper.tex",
)The bibliography result separates added, removed, modified, renamed, and unchanged entries. Rename detection needs a unique shared DOI, PMID, PMCID, arXiv ID, or ISBN. Citation comparison follows bounded literal includes and reports added, removed, and retargeted sites.
Visual diff planning keeps execution policy explicit:
from citeprobe import execute_visual_diff, plan_visual_diff
plan = plan_visual_diff(
"old-project",
"new-project",
"diff-work",
mode="source-only",
old_main="paper.tex",
new_main="paper.tex",
latexdiff_profile="safe",
old_number_mode="auto",
addition_color="teal",
deletion_color="magenta",
)
result = execute_visual_diff(plan)
for artifact in result.artifacts:
print(artifact.path, artifact.sha256)source-only runs latexdiff without TeX compilation. Native compilation
requires trust_project=True. Docker mode uses a no-network container with
resource limits and requires a digest-pinned image by default. Select
tex_engine="pdflatex", "xelatex", or "lualatex". safe_extract_zip()
provides bounded archive extraction for uploaded projects. Plans bind their
execution fields and commands to an integrity digest. Compilation can retry
without strike lines, and strikethrough_fallback=False disables that retry.
from citeprobe import (
load_benchmark_manifest,
manifest_template,
run_benchmark_manifest,
)
print(manifest_template())
manifest = load_benchmark_manifest("corpus.json")
outcome = run_benchmark_manifest(manifest)
print(outcome["passed"])Every artifact has an HTTPS URL, expected SHA-256, license declaration, backend, run count, and optional ground-truth count. The runner verifies the download before extraction and removes its temporary directory after each case. Default network requests reject local and private targets, resolve each redirect again, and pin the selected public address.
from citeprobe import render_results
markdown = render_results(results, "markdown")
sarif = render_results(results, "sarif")Supported values are available in REPORT_FORMATS. The renderer returns text
and does not write a file.
The UI is an optional FastAPI application:
from citeprobe.ui import create_app
app = create_app(
grobid_url="http://127.0.0.1:8070",
grobid_timeout=120,
)Install citeprobe[ui] first. The packaged command binds the app to
127.0.0.1; embedding code is responsible for its own server binding. The
GROBID URL is a server setting. A PDF upload can choose consolidation mode,
but it cannot supply a replacement URL.