A Python CLI tool for downloading Salesforce Attachment records and files using CSV-based record processing.
- Query Attachment records using Salesforce CLI authentication with WHERE clause filtering
- Process CSV files containing record IDs to download attachments
- In-memory pipeline — SOQL results stay in memory, no intermediate CSV files
- Per-batch download parallelism within each object
- Resume support — skip already-downloaded files on restart
- Atomic file writes (
.parttemp files +os.replace()) — no corrupted downloads - Graceful handling of OS errors (e.g. filenames exceeding filesystem limits)
- Execution reports (
report_downloaded.json,report_missing.json) with download URLs mise run verifyto check downloaded files exist on disk- Reuse sf CLI authentication (no separate OAuth setup required)
- Rich progress display with auto-detected renderer (Rich or tqdm)
- Advanced error handling with exponential backoff and connection pooling
- Flexible configuration via CLI arguments or .env file
- Intelligent filename collision detection using ParentId prefix
- Optional
--save-metadataflag to write SOQL result CSVs for audit/debug
This tool has been migrated from using Salesforce CLI (sf) subprocess calls to native simple-salesforce library operations for improved performance and reliability.
Before (v1.x):
- Used
sf data querysubprocess calls for SOQL queries - Used
sf data get-recordsubprocess calls for individual record retrieval - REST API downloads via
requestslibrary - Limited error handling and retry logic
After (v2.x):
- Native SOQL queries using
simple-salesforcelibrary - Direct REST API operations through
simple-salesforce - Advanced error handling with exponential backoff
- Connection pooling for optimal API usage
- Enhanced monitoring and health checks
- Performance: 2-3x faster execution through native API calls
- Reliability: Better error handling and automatic retries
- Monitoring: Comprehensive API usage tracking and health checks
- Scalability: Connection pooling supports higher throughput
- Maintenance: Reduced dependency on external CLI tools
- All existing CLI arguments and
.envconfiguration work unchanged - Output structure and file organization remain the same
- CSV file format requirements are identical
- Authentication still uses
sf org login(hybrid approach)
No action required! The migration is transparent to users. Simply:
- Update to the latest version
- Run
pip install -r requirements.txtto installsimple-salesforce - Use the tool as before - all existing workflows continue to work
If you encounter issues after updating:
- Authentication errors: Re-run
sf org login web --alias your-org - Import errors: Ensure
simple-salesforceis installed:pip install simple-salesforce - Performance issues: The new version should be faster; if not, use
--sync-onlyfor debugging - API limits: Monitor usage with
--debugflag for API call tracking
-
Salesforce CLI (
sfcommand)npm install -g @salesforce/cli
-
Authenticated Salesforce org
sf org login web --alias your-org
-
Python 3.8+
python3 --version
-
Python Dependencies
pip install -r requirements.txt
Key dependencies include:
simple-salesforce- Native Salesforce API client (replaces sf CLI subprocess calls)requests- HTTP client for REST API operationsrichortqdm- Progress display (optional)
-
Clone or download this repository
-
Install Python dependencies:
pip install -r requirements.txt
-
Optional: progress UI dependencies (already in
requirements.txt)richfor the rich terminal progress viewtqdmas a fallback renderer
The tool uses a CSV-based workflow where you provide CSV files containing record IDs (ParentIds), and it queries and downloads attachments for those records.
Basic usage:
python main.py --org your-org --records-dir ./records --output ./outputRequired arguments:
--org: Salesforce org alias--records-dir: Directory containing CSV files with record IDs
Optional arguments:
--output: Base output directory (default:./output)--batch-size: Number of ParentIds per SOQL query batch (default: 100). Download buckets are derived from this value (not separately configurable).--workers: Parallel workers for queries and downloads (default: 2, max: 8)--sync-only: Disable threading, run sequentially (default: disabled, use for debugging)--progress: Progress display mode (auto,on,off,tqdm)--verbose: Alias for default INFO logging--debug: Enable DEBUG console logging
Your CSV files in --records-dir must:
- Be UTF-8 encoded
- Have a header row with column names
- Contain an
Idcolumn with Salesforce record IDs (ParentIds) - Have at least one data row
Example CSV:
Id,Name,Description
001ABC123456789,Account 1,Main account
001ABC987654321,Account 2,Secondary account
aBo123456789ABC,Custom Record,Custom objectThe tool will:
- Read all CSV files from the records directory
- Extract record IDs from the
Idcolumn - Batch the IDs (default: 100 per batch)
- Query attachments with
WHERE ParentId IN ('id1', 'id2', ...) - Download attachment files organized by CSV filename
Step 1: Prepare CSV files Create CSV files with record IDs you want to process. Each CSV file will be processed separately, and attachments will be organized by the CSV filename.
Step 2: Run the tool
python main.py --org your-org --records-dir ./records --output ./outputWhat happens:
- Validates CSV files
- Extracts ParentIds from each CSV
- Queries attachments in batches using SOQL WHERE clause
- Downloads attachment files
- Saves metadata for reference
The tool uses parallel processing for both SOQL queries (Phase 2) and file downloads (Phase 3).
- Phase 2 (queries): All batches from all CSVs submitted to thread pool simultaneously
- Phase 3 (downloads): Objects processed sequentially, batches within each object downloaded in parallel
- Conservative defaults: 2 workers by default to avoid overwhelming the Salesforce API
- Connection pool: Each download worker borrows a connection from
SalesforceConnectionPoolviaget_connection()/return_connection()
--workersargument: Controls parallelism in both query and download phases (default: 2, max: 8)WORKERSenv var: SetWORKERS=4in.envfor persistent configuration--sync-onlyflag: Disable threading for debugging or API issuesSYNC_ONLYenv var: SetSYNC_ONLY=truein.envto disable threadingQUERY_TIMEOUT: Maximum time per individual batch query (default: 600 seconds / 10 minutes)
- Threading is enabled by default with 2 workers
- Each worker executes tasks sequentially (no internal parallelism within workers)
- Total parallelism equals the thread pool size
- Symmetric: same workers handle both queries AND downloads
- Start with default (2 workers) for safe operation
- For 10+ CSVs with many attachments: try 4 workers for better performance
- For Salesforce API rate limiting issues: try
--sync-onlyor reduce workers - Monitor logs with
--debugfor task timing and performance insights
# Default: 2 parallel workers
python main.py --org your-org --records-dir ./records
# Use 4 workers for faster processing
python main.py --org your-org --records-dir ./records --workers 4
# Disable threading for debugging
python main.py --org your-org --records-dir ./records --sync-only
# Set workers in .env file (persistent)
# In .env: WORKERS=4
python main.py --org your-org --records-dir ./recordsoutput/
├── report_downloaded.json # Execution report: successfully downloaded files
├── report_missing.json # Execution report: failed/missing files
└── csv_name_1/
├── metadata/ # (only with --save-metadata)
│ ├── batch_0_20260114_120000.csv
│ └── batch_1_20260114_120001.csv
├── a3xAAA111_invoice.pdf
├── a3xAAA111_receipt.pdf
└── a3xAAA222_contract.pdf
Each CSV file gets its own subfolder containing:
- Downloaded attachment binaries (directly in the object directory)
metadata/with per-batch SOQL result CSVs (only when--save-metadatais used)
Object directories are only created when there are files to download.
After each run, two JSON reports are written to output/:
report_downloaded.json— all successfully downloaded files with attachment_id, parent_id, filename, download URL, and file sizereport_missing.json— all files that failed to download with error details
On resume runs, reports are merged: new downloads are added, previously failed files that succeeded are moved from missing to downloaded.
Check that all downloaded files still exist on disk:
mise run verifyThis is useful when the output directory is synced to Google Drive or another service that may silently remove unsupported file types. Missing files are moved from report_downloaded.json to report_missing.json with error_type: "VerifyMissing".
Empty directories are cleaned up automatically after processing.
Filename Convention:
- Default format:
{ParentId}_{original_filename} - Example:
a3xAAA111_invoice.pdf
Collision Handling: When multiple attachments with the same name exist for the same ParentId, the tool automatically adds the Attachment ID:
- Format:
{ParentId}_{AttachmentId}_{original_filename} - Example:
a3xAAA111_00P1234_invoice.pdf
The tool supports loading configuration from a .env file in the project root directory.
Setup:
-
Copy the example file:
cp .env.example .env
-
Edit
.envwith your preferred values:# Salesforce org alias SF_ORG_ALIAS=your-org # Output directory OUTPUT_DIR=./output # Records directory RECORDS_DIR=./records # Log file path LOG_FILE=./logs/download.log # Batch size for SOQL queries BATCH_SIZE=100 # Console logging configuration VERBOSE=false DEBUG=false # Progress display configuration PROGRESS=auto WORKERS=2 SYNC_ONLY=false QUERY_TIMEOUT=600
-
IMPORTANT: Never commit the
.envfile to version control!
Supported Variables:
| Variable | Description | Default | CLI Override |
|---|---|---|---|
SF_ORG_ALIAS |
Salesforce org alias from sf CLI | None (use default org) | --org |
OUTPUT_DIR |
Base output directory | ./output |
--output |
RECORDS_DIR |
Directory containing CSV files | None (required) | --records-dir |
LOG_FILE |
Log file path | ./logs/download.log |
N/A |
BATCH_SIZE |
Number of ParentIds per query batch | 100 |
--batch-size |
WORKERS |
Parallel workers for queries and downloads | 2 |
--workers |
SYNC_ONLY |
Disable threading, run sequentially | false |
--sync-only |
QUERY_TIMEOUT |
Maximum time per individual batch query | 600 |
N/A |
VERBOSE |
Enable verbose console output (INFO level) | false |
--verbose |
DEBUG |
Enable debug console output (DEBUG level) | false |
--debug |
PROGRESS |
Progress display mode: auto, on, off, tqdm |
auto |
--progress |
Note: --save-metadata is CLI-only (no .env variable). It writes per-batch SOQL result CSVs to metadata/ subdirectories for audit/debug purposes.
Configuration Precedence:
- Command-line arguments (highest priority)
- Environment variables from
.envfile - Built-in defaults (lowest priority)
Control the number of ParentIds included in each SOQL query batch:
Via .env file:
BATCH_SIZE=100Via CLI:
python main.py --org your-org --records-dir ./records --batch-size 150Default: 100 record IDs per batch
Control parallel processing for both SOQL queries and file downloads:
--workers/WORKERS: number of parallel workers for queries and downloads--sync-only/SYNC_ONLY: disable threading for debugging or API issues
Via .env file:
WORKERS=4
SYNC_ONLY=falseVia CLI:
python main.py --org your-org --records-dir ./records --workers 4
python main.py --org your-org --records-dir ./records --sync-onlyDefault: 2 workers (parallel), threading enabled
Notes:
- Threading is enabled by default with conservative 2 workers
- Use
--sync-onlyfor debugging or when experiencing API rate limiting - More workers = faster processing but higher memory usage and API load
- Salesforce may rate-limit concurrent API calls; reduce workers if you see errors
The tool provides flexible logging with different verbosity levels:
Default (INFO level):
- Console shows main workflow progress, file download status, and results
- Log file contains all DEBUG details for troubleshooting
python main.py --org my-org --records-dir ./recordsVerbose mode (--verbose):
- Alias for default behavior (kept for compatibility)
- Shows INFO level logs on console
python main.py --org my-org --records-dir ./records --verboseDebug mode (--debug):
- Console shows all technical details: URLs, query previews, authentication details, etc.
- Useful for troubleshooting issues
python main.py --org my-org --records-dir ./records --debugVia .env file:
VERBOSE=false # Enable INFO level (same as default)
DEBUG=false # Enable DEBUG level with technical detailsVia CLI flags:
--verbose # Enable verbose output (INFO level)
--debug # Enable debug output (DEBUG level)The progress UI auto-selects a renderer (prefers rich, falls back to tqdm). You can force or disable it.
Via .env file:
PROGRESS=auto # auto, on, off, tqdmVia CLI flags:
--progress auto
--progress on
--progress off
--progress tqdmLogs are written to:
- Console: Configurable level (INFO by default, DEBUG with --debug)
- File:
./logs/download.log(always DEBUG level with full details)
Default output:
INFO - SALESFORCE ATTACHMENTS DOWNLOADER - CSV WORKFLOW
INFO - Found 1 CSV file(s): 20 records in 1 batch(es)
INFO - Batch 1/1: Querying 20 ParentId(s)
INFO - ✓ Query successful: 20 records
INFO - No filename collisions detected
INFO - Downloading 20 attachment(s)...
INFO - [1/20] invoice.pdf
INFO - ✓ Downloaded
INFO - [2/20] receipt.pdf
INFO - ⊙ Skipped (already exists)
INFO - Download complete: 1 downloaded, 19 skipped, 0 failed
INFO - WORKFLOW COMPLETE
Debug output adds:
- CSV file processing details
- SOQL query preview and length
- WHERE clause content
- Authentication details
- URL endpoints
- Bytes downloaded per file
The tool gracefully handles errors and provides clear error messages:
- Missing attachments (404 errors)
- Network failures (will stop the workflow)
- Invalid filenames
- Disk write errors
- Authentication expiry
- Invalid CSV files
- SOQL query length exceeded (with helpful suggestions to reduce batch size)
- Invalid SOQL syntax
- Insufficient permissions
SOQL query tasks use automatic retry logic (up to 3 attempts per batch). Download tasks use a single attempt (max_retries=1) with no timeout — failed downloads are counted and reported, and will be retried on the next run via resume.
Per-file OS errors (e.g. filename too long) are caught, logged, and recorded in report_missing.json with download URLs. Fatal authentication or network errors are re-raised and stop the workflow.
To avoid treating partial files as complete, downloads are written to a temporary folder first and only moved into the final destination on success.
- Temp folder:
output/.tmp_downloads/(global per--outputdirectory / org alias) - Lifecycle: the application cleans it before each CSV download phase and removes it after the phase completes
--org Salesforce org alias (required if not in .env)
--records-dir Directory containing CSV files with record IDs (REQUIRED)
--output Base output directory (default: ./output)
--batch-size Number of ParentIds per SOQL query batch (default: 100)
--workers Parallel workers for queries and downloads (default: 2, max: 8)
--sync-only Disable threading, run sequentially (default: disabled, use for debugging)
--save-metadata Write SOQL result batch CSVs to disk for audit/debug (default: off)
--progress Progress display mode: auto, on, off, tqdm
--verbose Enable verbose console output (INFO level)
--debug Enable debug console output (DEBUG level with technical details)
- Support limited to Attachment object (ContentDocument not yet supported)
- Batch size constrained by SOQL WHERE clause character limits
- Progress display requires
richortqdm(falls back to basic logging if unavailable)
Ensure you're logged in:
sf org display --target-org your-orgThe attachment may have been deleted. Check if ParentId still exists.
Ensure your sf CLI user has:
- Read access to Attachment object
- View All Data or appropriate object permissions
The query exceeds Salesforce's ~20,000 character limit. This happens when batch size is too large.
Solution:
python main.py --org your-org --records-dir ./records --batch-size 50Reduce --batch-size until the error disappears. Each Salesforce ID (18 chars) adds ~22 characters to the query.
You must provide a directory containing CSV files:
python main.py --org your-org --records-dir ./records- Try
--sync-onlyto verify it's not a threading issue - Check
--debuglogs for timing information - May indicate Salesforce API rate limiting (reduce workers or use --sync-only)
- Increase
QUERY_TIMEOUTenv var:QUERY_TIMEOUT=900(15 minutes) - Default is 600 seconds (10 minutes)
- Or use
--sync-onlyto eliminate timing pressure
- Reduce worker count:
--workers 2or--workers 1 - Each worker thread maintains state for active tasks
- More workers = more memory usage
- Reduce workers to decrease API load:
--workers 2 - Or use
--sync-onlyto eliminate parallelism entirely - Salesforce may rate-limit concurrent API calls
Basic usage with default batch size:
python main.py --org my-org --records-dir ./recordsCustom output directory:
python main.py --org production --records-dir ./records --output ./prod-attachmentsCustom batch size (process 200 IDs per query):
python main.py --org my-org --records-dir ./records --batch-size 200Using environment variables:
# Set in .env file
SF_ORG_ALIAS=my-org
RECORDS_DIR=./records
BATCH_SIZE=150
WORKERS=4
# Then run without arguments
python main.pysalesforce-attachment-download/
├── main.py # Main entry point
├── requirements.txt # Python dependencies
├── .gitignore # Git ignore file
├── .env.example # Environment configuration template
├── README.md # This file
├── src/
│ ├── __init__.py
│ ├── models.py # Data models (AttachmentRecord, BatchResult, ObjectQueryResult)
│ ├── exceptions.py # Custom exceptions
│ ├── utils.py # Logging utilities
│ ├── workflows/
│ │ ├── orchestrator.py # Three-phase workflow orchestration
│ │ ├── csv_coordinator.py # Phase 1: CSV discovery & processing
│ │ ├── query_coordinator.py # Phase 2: SOQL batch querying (in-memory results)
│ │ ├── download_coordinator.py # Phase 3: Per-batch parallel downloads
│ │ ├── thread_pool.py # Thread pool with retry logic
│ │ ├── error_handler.py # Centralized error handling
│ │ ├── directory_manager.py # Directory structure management
│ │ └── common.py # Shared utilities (ensure_directories)
│ ├── csv/
│ │ ├── processor.py # CSV file processing
│ │ ├── validator.py # CSV validation
│ │ └── enhanced_validator.py # Advanced CSV validation with field discovery
│ ├── query/
│ │ ├── executor.py # Query execution wrapper
│ │ ├── soql.py # Native SOQL execution via sf CLI (deprecated)
│ │ ├── soql_simple.py # Simple-salesforce SOQL queries
│ │ └── filters.py # WHERE clause building
│ ├── download/
│ │ ├── downloader_simple.py # Per-batch and single-file download functions
│ │ ├── filename.py # Filename sanitization and collision detection
│ │ ├── scan.py # Pre-download scan and skipped-files management
│ │ └── stats.py # Download statistics
│ ├── api/
│ │ ├── sf_auth.py # SF CLI authentication
│ │ ├── sf_client.py # REST API client (deprecated)
│ │ ├── sf_connection.py # Connection pool with instance_url
│ │ ├── sf_error_handler.py # Error handling with retries
│ │ ├── usage_monitor.py # API usage tracking
│ │ ├── sf_auth_adapter.py # Hybrid authentication adapter
│ │ └── field_discovery.py # Describe API field discovery
│ └── cli/
│ └── config.py # CLI argument parsing
├── records/ # CSV files with record IDs (user-provided)
├── output/ # Downloaded attachments (per-CSV subdirectories)
└── logs/
└── download.log # Execution logs
Potential improvements for future releases:
- ContentDocument/ContentVersion support - Handle newer Salesforce file storage
- Resume capability - Continue interrupted downloads from where they stopped
- Advanced filtering - Filter by date range, content type, file size, and other metadata
- Enhanced progress visualization - Improved progress indicators and real-time statistics
This project is licensed under the MIT License - see the LICENSE file for details.
Contributions are welcome! Please feel free to submit a Pull Request.