{
"summary": {
"fastp_version": "0.23.4",
"sequencing": "single end (75 cycles)",
"before_filtering": {
"total_reads":19014947,
"total_bases":1426121025,
"q20_bases":1368126463,
"q30_bases":1340057991,
"q20_rate":0.959334,
"q30_rate":0.939652,
"read1_mean_length":75,
"gc_content":0.501123
},
"after_filtering": {
"total_reads":10933431,
"total_bases":780019338,
"q20_bases":758743644,
"q30_bases":744983654,
"q20_rate":0.972724,
"q30_rate":0.955084,
"read1_mean_length":71,
"gc_content":0.498169
}
},
"filtering_result": {
"passed_filter_reads": 18724357,
"low_quality_reads": 1329,
"too_many_N_reads": 7,
"too_short_reads": 289254,
"too_long_reads": 0
},
$ fastp --detect_adapter_for_pe --overrepresentation_analysis --dedup --correction --cut_right --thread 10 --in1 fwd.fastq.gz --out1 clean/fwd.fastq.gz --unpaired1 clean/fwd.fastq.singletons.fastq --html stats_fastp.html --json stats_fastp.json
Here is the head of the file
stats_fastp.jsonfor a random single-end Illumina sequencing sample:after running it through fastp with the following command:
We can see that
after_filteringthere are 10'933'431 reads left in the cleaned FASTQ. However thefiltering_resultcategory tells us that as many as 18'724'357 passed the filter. This is a huge mismatch. What happened to the 8 or so million reads? Why did they get removed?