Assemble modelcif | Synthetic tests - #48
Draft
keiran-rowell-unsw wants to merge 63 commits into
Draft
keiran-rowell-unsw wants to merge 63 commits into
keiran-rowell-unsw wants to merge 63 commits into
Conversation
…which might be a superset
…d ModelArchive deposition
…CIF spec next push
…e true modelCIF entries
…CIF that has 20 residues to match metrics DUMMY
…ies, and snapshots pass
|
Warning Newer version of the nf-core template is available. Your pipeline is using an old version of the nf-core template: 4.0.3. For more documentation on how to update your pipeline, please see the Synchronisation documentation. |
|
Member
Author
- populate_modelcif.py: --seed accepts a rank_N-keyed *_seed.tsv (assets/DUMMY_SEED.tsv) or a single int; single-seed methods (alphafold2/colabfold/boltz via --random_seed) get a shared seed SoftwareParameter; multi-seed methods (alphafold3-style seed x sample filenames) get per-model seeds shown in model names, since the modelcif dumper registers only one QA metric type per class. - bin/utils.py: infer_model_seed() for AF3 'seed-N' and ColabFold '_seed_NNN' filenames. - Fix: modelcif System._before_write() only harvests parameter groups from system.software_groups; the container_image SoftwareWithParameters appended to system.software in 22aaf0b was silently dropped from the output mmCIF. - assemble_modelcif/main.nf: seed notes; stale TODOs refreshed.
…lcif.py - protocol: optional template_search_step (driven by software-details YAML) inserted upstream of the coevolution MSA step, with its own Data and shared step software. - per-program verified model facts (uses_templates, has_recycling, ColabFold has_relaxation) emitted as boolean SoftwareParameters; --param KEY=VALUE adds/overrides entries (ints, floats, bools coerced). - dummy software details: alphafold2 now carries a template_search_step. - CHANGELOG entry for the ASSEMBLE_MODELCIF feature batch. Co-Authored-By: pi <noreply@pi.dev>
…am keys - environment.yml: pin modelCIF to 1.7 (pip syntax, ==) so the conda-resolved versions match the recorded snapshots; regenerate the three versions.yml snapshot hashes accordingly. - populate_modelcif.py: --param use_templates now aliases to the canonical uses_templates key instead of adding a duplicate field. Co-Authored-By: pi <noreply@pi.dev>
TemplateSearchStep previously inherited the modeling Software, claiming
AlphaFold2 searched the templates. It now resolves a dedicated Software
from template_search_software.name (YAML) or --template_software
(the topic: val("template_search"), val(tool) emission), with version
from --template_version / YAML, drawn from the known search tools
(hhsearch/hhblits/jackhmmer/mmseqs2/AFDB). Falls back to the modeling
software group when the tool is unknown, preserving old behaviour.
- DUMMY_SOFTWARE_DETAILS: alphafold2 -> hhsearch, alphafold3 -> mmseqs2
(per upstream data pipelines), use_main_software false.
- new nf-test asserting the attribution.
Co-Authored-By: pi <noreply@pi.dev>
…Biology-Computing/proteinfold into assemble_modelcif
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part implements an ASSEMBLE_MODELCIF{} process as sketched out https://github.com/orgs/nf-core/projects/151
assemble_modelcif_X has ext.args config files to test creation of .bcif, or PAE embedded or linked
Test data: the following .cif files and dictionaries are used for testing, but not included in the PR to avoid binaries and line count inflation, they may be included in in test-datasets/proteinfold
mmcif_ddl.sdbmmcif_ma.sdbmmcif_pdbx_v50.sdbma-atiya1-gk-01.cifDatabases: databases will not be handled in .data.Datagroup.ReferenceDatabasein this PR. populate_modelcif.py is getting quite long already. Plus, it's a separate concept that can tie into the work done for reference dataset at NCI.
nf-core#575 might make this database handling easier, if considered valuable