Skip to content

Latest commit

 

History

History
111 lines (99 loc) · 8.25 KB

File metadata and controls

111 lines (99 loc) · 8.25 KB

DeepStrain — Roadmap (forward-looking)

Complements CLAUDE.md (current status per sub-project) and JOURNAL.md (dated history). This file captures the next high-leverage moves and the standing guardrails. Each item says what, why it's high-leverage, which sub-project / where the detail lives, and status. Added 2026-06-20 from a cross-project review ("legs" analysis).

All three arcs are currently PARKED with honest results (two wins, one modest, one honest negative). These items are the cleanest ways to strengthen what we already have on the same data, plus the one guardrail on a parked thread.


P1 — Echo non-detections → real UPPER LIMITS (highest leverage)

  • What: add an injection-efficiency curve to the echo comb search — i.e. measure, per echo spacing Δt (≡ λ), the amplitude you would have detected. That converts the current honest "non-detection" into a quantitative "we exclude λ above X" exclusion.
  • Why high-leverage: same data, but a genuinely stronger and more publishable result — an exclusion is a constraint, a non-detection is not. Flagged honestly in the leg-8 review.
  • Where: echoes/. Detail + the v1/v5 sensitivity machinery to extend: echoes/notes/lab_notebook.md. (v1 already has a sensitivity-curve harness 06; this generalizes it to a per-Δt efficiency → λ map.)
  • Status:DONE (2026-06-20, v6). scripts/11_upper_limits.py: per-Δt exclusion curve at N=300. GW150914: exclude amplitude ≥ A90=1.65σ at predicted Δt (A50 1.33σ); GW151226: ≥1.55σ at its canonical Δt=0.0579 s. Smooth across all spacings; stress-tested (statistic verified, threshold not glitch-driven). (γ=0.7 fixed.) Update (E1, 2026-06-25): the ML scorer does NOT tighten this — through the honest production path A90 ML≈comb (0.98×); v5's ~1.2× edge is a 50%-point effect that vanishes at the 90%-exclusion level. The comb UL stands. See PLAN.md E1 / echoes lab notebook.

P1 — Multi-event no-hair δ STACKING

  • What: combine the no-hair deviation δ across multiple events, not just the single spine event (GW250114). Single-event σ(δ) ≈ 0.24; stacking is the clean way to sharpen it.
  • Why high-leverage: the one place more data directly tightens a real GR test. Already noted as a v2 direction; the amortized SBI network is built and calibrated, so this is mostly a hierarchical-combination layer on top of proven infra.
  • Where: ringdown_spectroscopy/ (no-hair arc, v2/v3 — COMPLETE & calibrated). Detail: ringdown_spectroscopy/notes/lab_notebook.md.
  • Status: ⚠️ METHOD ✅ / real payoff ❌ PARKED (2026-06-20, v5 + stress-test). 12_stacking.py validated the stacking METHOD — σ(δ) tightens as √N on informative injections (N=8 → 0.095 vs ideal 0.097, unbiased, calibrated), gated. BUT the north-star stress-test (13_more_events.py) showed only GW250114 actually measures δ; all 7 fainter public events return ≈ the prior. So the v5 "GW250114+GW150914 → 1.3× tighter" was a Gaussian-approx-of-prior artifact (corrected) — there is effectively ONE informative real event. Real multi-event sharpening is blocked by the per-event SNR information wall (only SNR~80-class events measure δ). Come-back-later = more very-loud events, or an NPE that extracts δ at lower SNR (likely information-limited, like tone-count). v6 (2026-06-20) MAPPED THE WALL: 14_delta_threshold.py swept injected ringdown loudness and measured σ(δ) vs SNR — δ only becomes informative (σ/prior < 0.90) at ringdown SNR ≳ 37, and even at the top of the NPE's trained loudness it's just ~13% tighter than the prior; GW250114 (real, σ/prior 0.83) sits right at that edge. So the stacking starvation is now quantitative, not anecdotal: every public event lands at-or-below the informative threshold. Seed-robust, gated.

P2 — Higher-N injection campaigns where claims are UNDERPOWERED

  • What: re-run the underpowered claims at N ≈ 300–500 injections. Specifically the leg-8b "sensitivity reversal" rested on N = 25 (confidence intervals overlap, so it's not yet real or refuted).
  • Why: cheap and decisive — settles whether the effect is real instead of leaving it ambiguous. Low effort, high clarity.
  • Where: echoes/. Detail: echoes/notes/lab_notebook.md.
  • Status:DONE (2026-06-20). (a) Upper limits run at N=300. (b) The specific leg-8b "sensitivity reversal" SETTLED: re-ran 08 --n-trials 300 → the in-band family differences are REAL & physically sensible (f0=320/γ=0.9 genuinely easier, +6–7σ; f0=150/γ=0.5 harder), NOT a pathology — the N=30 overlap was just underpower. The one true anomaly (out-of-band control not collapsing) is the known whitened-domain artifact (valid v4 raw test = 10%). No pathological reversal survives.

GUARDRAIL — Do NOT throw more ML at the tone-count gap

  • What: keep the ringdown v4 tone-count thread PARKED. Do not iterate more classifier architectures on the same data.
  • Why: the gap is information-limited, not legibility-limited — independently confirmed (leg 2 and leg 7, and our own six-attempt diagnostic chain ending in a calibrated-but-weak AUC ~0.61). The real lever is more SNR / a coherent multi- detector model, not a fancier net on the same parked data.
  • Where: ringdown_spectroscopy/ v4 (PARKED — honest negative). Detail + the full six-attempt table: ringdown_spectroscopy/notes/lab_notebook.md.
  • Status: 🅿️ PARKED intentionally. Revisit only with more data / a coherent model / multi-event stacking / explicit Bayesian model selection.

LONG-HORIZON PROJECTS (L1–L7) — tracked, never dropped for size

Full scoping in RELATED_WORK.md (§ Long-horizon projects), added 2026-08-15 from a literature sweep. Standing rule: effort is not a reason to drop an item. There is no deadline; the Mac runs unattended; long-running jobs are normal. An item leaves the list when it is done or measured to be impossible — never because it looked big.

project effort buys
L1 cheap-template dense bank (de-chirping, ~8× per core → ~13k templates) weeks pushes the one axis Follow-up A named as the dominant loss
L2 deep background → 1/century (~353 segs ≈ 1000 yr, ~14 h unattended) days a deeper rung and ±33% → ~±18% on the existing one
L3 orthonormal-mode adoption across the ringdown arc (retrain NPE, recalibrate) weeks could overturn the parked v4 tone-count negative
L4 coherent network echo search (phase-marginalized, cf. arXiv:2512.24730) weeks field-standard echo sensitivity; nulls become comparable
L5 three-tone spectroscopy (GWTC-5.0 reports the first 3-tone measurement) read, then weeks extends the no-hair test beyond 220+221
L6 larger unlabeled pool for the SSL backbone days tests whether N4's data-wall win keeps growing
L7 S251112cm (FAR 1/6.2 yr subsolar candidate) blocked: O4c not public. o4c_release_watch.py is the trigger

Known blockers carried forward (context for the above)

  • PBH subsolar: template-bank density wall — subsolar needs ≤0.1% Mc spacing (~1,600+ templates); 1,619 was our laptop ceiling. ⚠️ "Intractable locally" NO LONGER HOLDS (2026-08-15): the field's ratio-filter de-chirping reports ~8× per-core speedup, i.e. ~13k templates on the same hardware — a real move down the density sweep toward the 0.72 oracle. Reclassified from blocker to L1, a tracked long-horizon project. Do not restate it as intractable. (pbh v2 PARKED; coincidence win +1.37× stands.) Build C DONE (2026-06-20, L4 VM): the "lower FAR needs more data" item is closed — coincidence is FAR-robust (graceful to 1/year; @1/day reproduces the +1.37×; @1/year still beats single-det floor ~1.2×). See primordial_blackhole_search/RESULTS.md.
  • Ringdown tone-count: information-limited (see guardrail above).