Observation-conditioned selection-aware conformal triage for time-domain astronomy
Bishal Neupane
PAPER · v1.0 · 2026-09-14 · human
Abstract
The difficulty facing machine-assisted astronomy is no longer the construction of another high-capacity predictor. The harder problem is deciding whether an output remains scientifically defensible when the observation process, population, measurement precision and available follow-up all change at once. We develop an observation-conditioned selection-aware conformal triage procedure, OCSAT, for sparse astronomical alerts. The method has four pieces. It estimates the latent population after correcting the labelled sample for its detection and follow-up probability. It performs inference with the physical observation operator that generated the data rather than replacing that operator with a learned feature map. It evaluates model discrepancy on observations that were not used to fit the current physical state. It then calibrates the final action threshold online against a declared follow-up budget. We audit the independent literature around selection bias, measurement error, conformal inference, active follow-up and physics-informed learning and find that these components are established separately, but we did not find a directly matching framework in which selection correction, observation-conditioned physical residuals and drift-aware conformal action control are joined at the upstream astronomical decision boundary. The empirical study combines current public archive statistics with a controlled survey simulation based on a Bazin transient model, heteroscedastic photometric errors and changing cadence. In a 50,000-object latent population, inverse-selection weighting reduces absolute error in the three population means from 0.505, 0.169 and 0.204 to 0.054, 0.024 and 0.040, and reduces the Frobenius covariance error from 0.227 to 0.117. Under a four-state change in measurement quality, an adaptive threshold keeps the corrected follow-up fraction near the one percent operating point, whereas a fixed threshold becomes increasingly conservative. A separate held-out injection experiment shows how an explicit physical residual responds to an unmodelled fast component and quantifies the loss in sensitivity as photometric precision degrades. These results support a narrower claim than universal superiority. The main gain comes from changing what is calibrated. The calibrated object is not only a prediction score, but a scientific action conditioned on how the measurement was made.