Skip to content

Match writer's best-sample metric to analyze's - #253

Open
fnachon wants to merge 9 commits into
HannesStark:mainfrom
fnachon:fix/writer-analyze-best-sample-metric
Open

Match writer's best-sample metric to analyze's#253
fnachon wants to merge 9 commits into
HannesStark:mainfrom
fnachon:fix/writer-analyze-best-sample-metric

Conversation

@fnachon

@fnachon fnachon commented Jul 10, 2026

Copy link
Copy Markdown

Fixes #127.

FoldingWriter.write_on_batch_end() selected the best sample of a batch using 0.8*iptm + 0.2*ptm, while analyze_utils.get_best_folding_sample() selects using 0.8*design_to_target_iptm + 0.2*design_ptm. Since these are generally different values, the refolded structure written to disk by the writer was often not the same sample that the analysis/filtering step scored and reported as best. Changed the writer to use the same metric as analyze.

fnachon and others added 9 commits January 10, 2026 15:54
Changes made to run without errors on the Mac MPS device: torch.autocast, number of devices and workers to use on M1-5 chips, workaround for CUDA-specific code, handling of float64 incompatibilities for MPS.
Replace hardcoded torch.autocast("cuda") with device-agnostic
device_type=tensor.device.type in confidence_utils, inverse_fold,
and writer modules introduced in the upstream merge.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Python pickle does not preserve RDKit atom-level SetProp values. When
PyTorch DataLoader spawns worker processes (default num_workers=1 on
macOS), self.canonicals is pickled and all atom 'name' properties are
lost, causing KeyError in process_atom_features.

Fix: load all required molecules directly from the moldir zip inside
each get_sample() / get_feat() call instead of using the pickled
self.canonicals. The moldir zip handle is cached per-process by
_get_zipfile(), so there is no repeated I/O overhead.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…ders

- Disable pin_memory on MPS (unsupported, causes UserWarning)
- Enable persistent_workers when num_workers > 0 (avoids repeated
  worker init overhead and the PL suggestion warning)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
FoldingWriter.write_on_batch_end() selected the best sample using
0.8*iptm + 0.2*ptm, while analyze_utils.get_best_folding_sample()
uses 0.8*design_to_target_iptm + 0.2*design_ptm. The mismatch meant
the refolded structure written to disk was often not the same sample
that analysis/filtering scored and selected as best.

Fixes HannesStark#127

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Best sample selection uses different confidence metrics in writer vs. analyze

1 participant