dataprof profiles CSV, JSON, Parquet, DataFrames and Arrow tables in one call: types, nulls, distributions, patterns and a quality score. Then it can gate your pipeline on what it found. It runs locally, gives the same numbers for the same data whichever way you load it, and says "can't tell" instead of guessing.
A Rust library with a Python package on top. No server, no account, no network.
pip install dataprof # or: uv pip install dataprofWheels for CPython 3.10 to 3.14 on Linux, macOS and Windows, with no Python dependencies.
import dataprof as dp
report = dp.profile("orders.csv") # also .json, .jsonl, .parquet, DataFrames, dicts
print(report.rows, "rows,", report.columns, "columns, quality", report.quality_score)
amount = report["amount"]
print(amount.data_type, amount.mean, amount.null_percentage)Ask it what deserves attention. Each finding has a stable code and the evidence behind it, never a raw value:
for finding in report.findings():
print(finding.severity, finding.code, finding.column)
# warning locale_numbers price ("10,50"-style numbers, left out of the stats)
# warning null_heavy email
# info sensitive_pattern emailState what "good enough" means and get a verdict: pass, fail, or inconclusive when the data can't prove it either way.
result = report.check(min_quality_score=90, max_null_percentage={"customer_id": 0, "*": 20})
if not result.passed:
for check in result.violations:
print(check.code, check.column, check.message)Or from CI, with no code at all:
python -m dataprof.check orders.csv --min-quality 90 --max-null "*=20"
# exit 0 = pass, 1 = fail, 2 = inconclusive or bad inputA token-bounded summary for an LLM. Values that look like personal data are never echoed:
print(report.to_llm_context(max_tokens=500))cargo add dataprofuse dataprof::{Profiler, QualityPolicy, Verdict};
let report = Profiler::new().analyze_file("orders.csv")?;
let result = QualityPolicy::new().min_quality_score(90.0).evaluate(&report)?;
assert_eq!(result.verdict, Verdict::Pass);Minimum supported Rust: 1.96.
- Same data, same numbers. CSV, Parquet, pandas, polars and Arrow of the same values produce the same profile, on every engine. CI checks it.
- Bounded memory. Files larger than RAM stream through fixed-size accumulators.
- Honest verdicts. A sampled or partial scan never passes a claim about the whole file; scores carry their confidence interval.
- Absence is not zero. A metric that wasn't computed is
None, never a plausible default. - Private by default. Reports and summaries carry counts and patterns, not your values.
Beta, and moving fast toward a stable 1.0 contract. Some things are still narrow, and we'd rather tell you than have you find out: numbers written with a decimal comma are detected and flagged but not yet parsed, and markers like NA or N/A aren't treated as nulls yet. The release notes list what changed and what's known.
Found something wrong? Open an issue. A file that profiles badly is the most useful bug report there is.
- Why dataprof?: an honest comparison with pandas, polars and ydata-profiling, including when to pick them instead
- Getting started: the metrics and how to read them
- Python API: every function, the gate and the CI entrypoint
- Examples: a messy CSV, an ETL gate, before and after cleaning, runnable from a clean checkout
- Agent workflows: AGENTS.md, Cursor rules and Claude Code skills
- Report schema: the saved document and its guarantees
- Database connectors and feature flags: Rust-side options
- Changelog and contributing
Use Cite this repository in the GitHub sidebar, which builds APA and BibTeX from CITATION.cff. The citation is for the software; benchmark material lives in scalcom2026-dataprof.
Either the MIT License or the Apache License, Version 2.0, at your option.
