Outputs new JSON file with dupes from original removed, as well as logging which records and fields changed
To run the program, open a terminal and then:
java -jar java-json-deduplication.jar your_file.json
Then find the deduplicated.json and deduplication_log_[datetime].log file upon completion
The requirements didn't state what should happen under circumstances like the duplicate having a value and the preferred record not having a value. Should the non-empty value be preserved or discarded? I chose to remove the duplicate wholesale. If this were a real work assignment, I'd ask the PM/product owner/etc. what the client would prefer.
- Data from newest-date entry preferred
- Dupe _id and dupe email fields count as dupes, both must be unique. Dupes elsewhere don't matter.
- If a dupe is found and both records have identical entryDate, entry last in the list is preferred
- No need to worry about large files, program allowed to do everything in memory
- Provide log of changes including: source record, output record, individual field changes (from/to)