stat updates on arXiv.org
stat updates on the arXiv.org e-print archive.
A Multi-Cohort Validation of Censoring-Aware Conformal Lower Predictive Bounds for Pathology Survival Models
oai:arXiv.org:2608.04025v1
arXiv:2608.04025v1 Announce Type: new
Abstract: Whole-slide survival models commonly provide risk rankings without calibrated statements about individual event times. We evaluate fixed-cutoff drcosarc, a post-hoc conformal wrapper for discrete-time multiple-instance learning survival heads using frozen UNI2-h representations, in an internal 18-configuration sweep across five TCGA cohorts and an external five-configuration evaluation across three CPTAC cohorts. We distinguish configuration--fold--split summaries of the inverse-probability-of-censoring-weighted (IPCW) estimate and median lower predictive bound (LPB) from a hierarchy-aware patient-ensemble estimand of the mean drcosarc--naive LPB difference. At $\alpha=0.1$, the drcosarc IPCW estimate was nearest 0.90 in KIRC, LUAD, and STAD. Patient-ensemble drcosarc--naive intervals excluded zero in KIRC, KIRP, STAD, UCEC, and CPTAC-CCRCC, but included zero in internal LUAD, CPTAC-LUAD, CPTAC-UCEC, and the internal LUSC extension. In a 20-replicate low-censoring semi-synthetic setting with known event times, drcosarc empirical coverage was 0.9129 [0.9053, 0.9207]. An exploratory analysis supported a head-error-by-censoring interaction within that data-generating process. In a two-cohort ABMIL sensitivity analysis, increasing the hazard grid to $K=16$ raised localized marginal IPCW estimates above the prespecified 0.87 threshold and yielded positive paired LPB differences, although worst-group estimates remained below 0.87. Overall, performance was cohort dependent, and its interpretation changed with the patient-level unit, estimand, and censoring assumptions.
Statistical learning theory and Occam's razor: Regularization
oai:arXiv.org:2608.04049v1
arXiv:2608.04049v1 Announce Type: new
Abstract: The principle of Occam's razor, which instructs us to prefer simplicity in inductive inference, has attracted much scrutiny both in the philosophy of science and in machine learning. In either field, however, a justification for the principle has been elusive. In this paper, building on an earlier "core argument," I spell out a justification from statistical learning theory for the procedure of regularization: for trading off fit for simplicity. The means-ends argument is that in order to profit from theoretical reliability and "what-you-see-is-what-you-get" guarantees, one must implement a certain preference for simplicity over fit. This is a genuine methodological justification, which neither collapses to a purely pragmatic principle that we prefer simplicity for its own sake, nor to an ontological assumption that the truth is simple.
Discrete-Time Survival Analysis for Heart Failure Mortality Prediction
oai:arXiv.org:2608.04140v1
arXiv:2608.04140v1 Announce Type: new
Abstract: Accurate heart-failure prognosis relies on tracking clinical risk over time, yet many machine-learning applications mishandle right-censored survival data by either discarding a patient's observation time or using it as a predictor. Discarding time ignores survival context, while using follow-up time as an input feature introduces severe target leakage that inflates apparent accuracy. We address this by proposing a discrete-time person-period framework for heart-failure mortality classification. Using the UCI Heart Failure Clinical Records cohort ($n=299$, 96 deaths), we transform the data into interval-level binary outcomes and benchmark a Cox proportional hazards baseline against person-period complementary log-log GLM and GAM models, alongside person-period random forest, XGBoost, random survival forest, and DeepSurv classifiers. The person-period GLM reproduces the Cox hazard ratios and concordance, validating the transformation, while the GAM captures significant nonlinear predictor effects and provides the best balance of discrimination and generalization; the flexible classifiers achieve strong raw performance but overfit. Finally, we quantify the leakage effect directly, including observed follow-up duration raises classification AUC from roughly 0.73 to nearly 1.00, confirming that follow-up duration must not be used as a baseline predictor. Overall, these results establish a survival-aware framework that combines flexible classification with valid time-to-event structure.
Regression-Based Proximal Reconciliation of Conflicting Trials with Unmeasured Effect Modifiers
oai:arXiv.org:2608.04202v1
arXiv:2608.04202v1 Announce Type: new
Abstract: Randomized controlled trials with similar protocols may yield conflicting findings when the distribution of relevant effect modifiers differs across study populations. Yet no formal statistical framework exists for defining and assessing whether conflicting trials are reconcilable, despite the importance of this question for evidence synthesis and regulatory decision making. To address this gap, we develop a causal inference framework for evaluating conditional and marginal reconcilability on additive and multiplicative scales in the presence of unmeasured effect modifiers. Within this framework, we use proxy variables for hypothesized unmeasured effect modifiers to develop regression-based tests of conditional reconcilability under parametric structural models. To assess marginal reconcilability, we extend existing transportability methods and develop an equivalence testing framework. We also introduce a reconciliation proportion to quantify the degree of marginal reconciliation. We illustrate these methods using the conflicting Meis and PROLONG trials of 17-alpha-hydroxyprogesterone caproate for preventing recurrent preterm birth. The analyses provided limited evidence that unmeasured effect modifiers such as cervical length, as captured by the selected proxies, were sufficient to marginally reconcile the trials. These findings demonstrate how proximal reconciliation methods may help regulators, researchers, and clinicians evaluate whether differences in study populations explain conflicting trial findings.
Estimating Heterogeneity in Travel Mode Choice Shifts with Causal Forests
oai:arXiv.org:2608.04208v1
arXiv:2608.04208v1 Announce Type: new
Abstract: Objectives: While causal analysis of travel behavior is an emerging field, estimating heterogeneity in mode choice through causal modeling remains unexplored. This study demonstrates the application of a novel causal method, causal forest, to quantify the heterogeneity in travel mode choice shifts caused by the COVID-19 pandemic.
Methods: We applied causal forests, a non-parametric causal machine learning method, to 802,935 trip records from the 2017 and 2022 waves of the National Household Travel Survey. The 2017 wave serves as the pre-pandemic control group, while the 2022 wave represents the treatment condition. Within the potential outcomes framework, we estimate average treatment effects (ATE), heterogeneous treatment effects (HTE), and conditional average treatment effects (CATE) across diverse socio-demographic groups and trip characteristics.
Findings: Our results reveal an estimated ATE of a 1.86 percentage point (pp) increase in car-mode share, contrasted with decreases of 0.38 pp and 1.57 pp in public transit and walking, respectively. The largest increases in car use appeared for short-distance trips (one mile or less), households with annual incomes exceeding USD 200,000, and female travelers.
Novelty: This is one of the first applications of causal forests to travel mode choice, and the first to use causal machine learning to estimate the pandemic's causal effect on mode choice analysis.
Practical Applications: This study discusses methodological advantages, inherent assumptions, and limitations of causal forests within the context of transportation planning. This methodology is applied to COVID-19 travel data to illustrate how causal heterogeneity analysis can offer a deeper understanding of changes in mode choice. These insights are valuable for planners and policymakers in making policies related to mode shifts under an intervention.
Multimodal Alignment Through Joint Kernel Entropic Gromov--Wasserstein Optimal Transport
oai:arXiv.org:2608.04234v1
arXiv:2608.04234v1 Announce Type: new
Abstract: We study the problem of aligning data from multiple modalities into a shared representation space, focusing on settings where strong pretrained unimodal encoders are available but cross-modal paired data are scarce. We propose a structure-preserving alignment framework, joint kernel entropic Gromov--Wasserstein Optimal Transport (JK-EGW), which maps multiple modalities into a common latent space by minimizing a quadratic optimal transport objective. JK-EGW leverages fine-grained similarity relationships within and across modalities to construct a global affinity kernel instead of relying on raw feature-space distances. Our framework naturally provides explicit control over the geometry and distribution of the latent embedding. On the theory side, we establish parametric sample complexity rate of $n^{-1/2}$, matching the corresponding rates for standard, entropic and Gromov--Wasserstein optimal transport. On the algorithmic side, we derive a scalable alternating procedure to solve JK-EGW with entropic optimal transport (EOT) updates through a low-rank kernel approximation and a variational lifting. This lifting scheme effectively relieves the burden of a quadratic objective, and allowing us to take the advantage of existing EOT solvers. Empirically, we focus on post-hoc alignment of embeddings from pretrained encoders in data-scarce regimes, and show that our proposed method achieves improved multimodal retrieval performance compared to existing alignment baselines.
When Is a Conformal Guarantee Fair? Auditing Silent Subgroup Under-Coverage in Alzheimer's Disease Longitudinal Prediction
oai:arXiv.org:2608.04254v1
arXiv:2608.04254v1 Announce Type: new
Abstract: Longitudinal prediction of Alzheimer's disease biomarkers increasingly informs clinical decisions, and a forecast is only useful if it also reports how much to trust it. Conformal prediction supplies this by wrapping any forecaster in a prediction band with a finite-sample coverage guarantee under exchangeability. However, standard population-level conformal prediction guarantees only marginal coverage and may mask substantial under-coverage within clinically important subgroups. We introduce a general mechanism-driven framework for auditing and repairing such subgroup under-coverage. Across two cohorts (ADNI, OASIS-3), two base forecasters, and nine attributes spanning genetic risk, demographics, and clinical severity, we find that population-level bands under-cover high-risk subgroups in 57 of 68 audited combinations, despite achieving nominal marginal coverage. We trace these failures to two mechanisms: (A) \emph{rarity}, where a group-conditional band calibrated on only $n$ patients covers at most $k/(n+1)$; and (B) \emph{tail-heaviness}, where a population-wide band is too narrow for a heavy-tailed subgroup and additional data cannot close the gap. Under-coverage falls disproportionately on patients with high genetic risk and disease severity (6.1 pp mean deficit, 95\% CI [3.3, 8.9]), while demographic groups remain at the target level on average (0.0 pp, CI [$-1.9$, 1.7]). We pair each mechanism with a corresponding conformal correction: cross-conformal pooling for rarity, per-subgroup calibration for tail-heaviness, and a coverage-safe marginal floor when both arise. Together, these corrections restore target coverage for nearly every high-risk subgroup across both cohorts and forecasters.
Diffeomorphic Markov Chain Monte Carlo: fast mixing for heavy-tailed distributions
oai:arXiv.org:2608.04284v1
arXiv:2608.04284v1 Announce Type: new
Abstract: We introduce a new class of uniformly ergodic MCMC algorithms, termed Diffeomorphic Contraction Sampler (DCS), and provide fast non-asymptotic mixing guarantees for DCS targeting distributions on $\R^d$ with arbitrarily heavy polynomial tails. DCS provides a solution to a well-known problem for MCMC samplers, which typically struggle with the combination of unbounded high-dimensional state space and vanishing gradients.
The DCS pulls back a target on $\R^d$ onto a Euclidean ball $B(R)\subset\R^d$ and then samples from the transformed density on the convex set $B(R)$ via algorithms such as the Ball Walk, Hit-and-Run and others. A radial diffeomorphic contraction is chosen so that the pull-back density on $B(R)$ is bounded, implying uniform ergodicity for \textit{all} targets with a finite polynomial moment. Non-asymptotic bounds for DCS require stronger assumptions such as log-concavity of the pull-back density. In practice, this is achieved approximately by a preconditioned automorphism of the ball $B(R)$, tuned via Variational Inference.
Numerical simulation tests demonstrate that the DCS outperforms significantly the No-U-Turns sampler on multi-dimensional heavy-tailed targets arising as real-world posteriors in PosteriorDB benchmark. DCS also numerically outperforms in high-dimensional examples recently developed spherical projection samplers for heavy-tailed target distributions.
Fixed-Point Characterisations of Extremal Distributions under Partial Distributional Constraints
oai:arXiv.org:2608.04315v1
arXiv:2608.04315v1 Announce Type: new
Abstract: We present a methodological framework for solving robust inference problems with partially specified distributions over measurable subsets of a parameter space. Partial specifications define a set of admissible distributions. The goal is to determine extremal values (over these admissible distributions) for statistical quantities, where these quantities---these objective functions---are ratios of expectations of analytic functions, continuous functions, piecewise continuous functions, and uniform limits of piecewise continuous functions. We show that extremal values are approached by sequences of admissible distributions, whose limiting extremal distributions are characterised by fixed-point conditions on their support locations. This characterises where extremal distributions place probability mass and yields a practical computational framework for solving the corresponding optimisation problems. We establish convergence and asymptotic properties of the resulting extremal distributions and extremal objective function values. This work extends robust inference methods (e.g. robust Bayesian inference) by combining extremal-distribution reduction, fixed-point characterisation, and approximation-based analysis within a unified framework.
Bivariate Prior Specification for Bayesian Decision Making in Early Phase Clinical Trials
oai:arXiv.org:2608.04335v1
arXiv:2608.04335v1 Announce Type: new
Abstract: Bayesian Go/No-Go decisions with co-primary endpoints require specifying prior distributions under the Normal-Inverse-Wishart framework; however guidance on how prior hyperparameters influence trial decisions remains limited. We propose a calibrated prior specification framework for bivariate Go/No-Go decisions. Skeptical and enthusiastic priors are calibrated so that each assigns a target probability to a clinically relevant decision region. We prove that for any prior precision $\kappa > 0$, a unique scale parameter $\lambda_0$ achieves the target calibration. Operating characteristics are evaluated across different $\kappa $ via simulation and applied to a phase~3 telitacicept lupus trial.The simulation result indicates $\kappa$ is the primary driver of prior discrimination. At $\kappa = 1$, the go rate difference between priors was 0.07; at $\kappa = 10$ it reached 0.56, with false positive rates below 0.01. Operating characteristics were robust to the degrees of freedom parameter $\nu_0$ and prior correlation $\rho_0$, supporting a default of $\nu_0 = 2$. In the lupus application, prior sensitivity was negligible at $\kappa = 1$ but at $\kappa = 10$ the enthusiastic go rate was three times the skeptical rate at small sample sizes. The framework reduces prior specification to two choices: the prior center and the prior precision $\kappa$. The identification of $\kappa$ as the dominant parameter, together with the cautious choice of $\kappa$ before the trial, motivates adaptive approaches to prior precision.
Best for which estimand? A known-truth benchmark of longitudinal-matching and target-trial-emulation methods for time-varying treatments
oai:arXiv.org:2608.04414v1
arXiv:2608.04414v1 Announce Type: new
Abstract: On a non-collapsible survival mechanism, longitudinal-matching and target-trial-emulation methods are not competing estimators of one truth but answers to different causal questions, so a benchmark that scores them against a single "true hazard ratio" fabricates bias. We provide the direct comparison of relative efficiency, variance estimation, and model sensitivity that reviews find lacking. On a deliberately non-collapsible continuous-time Cox mechanism with known truth, the dominant families (sequential Cox, sequential stratification, risk-set matching, and inverse-probability-of-treatment-weighted (IPTW) marginal structural models) target numerically distinct causal estimands (marginal, conditional, two average-treatment-effect-on-the-treated, and intention-to-treat versus per-protocol). First, we quantify the phantom bias a shared marginal truth fabricates: 0.32-0.33 log-cumulative-hazard-ratio units for the matching estimators and 0.15 for the conditional method; the associational naive time-dependent Cox sits 0.76 away, a total discrepancy compounding the estimand gap with confounding. Second, a rank reversal: the recommended method flips with the target estimand, and a low-variance off-target estimator can still win on mean-squared error. Third, a cross-family variance result: the cluster-robust sandwich is closer to nominal for the trial-stacking estimator (0.90) but under-covers the matching estimators (0.77-0.82), which a prespecified n=500 bootstrap sub-study brings to 0.95-0.96. Fourth, model sensitivity: omitting a confounder induces 0.45-0.50 log-hazard-ratio bias and undercoverage, and intention-to-treat and per-protocol effects diverge as switching increases; a heart-transplant analysis illustrates these. On a second mechanism three of four findings replicate, the rank reversal attenuating and model sensitivity proving calibration-dependent.
Semiparametric robust mixture of experts based on nonparametric maximum likelihood
oai:arXiv.org:2608.04561v1
arXiv:2608.04561v1 Announce Type: new
Abstract: The mixture of experts (MoE) model provides a flexible approach for modeling heterogeneous regression relationships by allowing covariate-dependent mixing through a gating network, but most existing MoE models rely on parametric assumptions for expert error distributions, typically Gaussian, which can lead to inefficiency and sensitivity to outliers or heavy-tailed behavior when misspecified. We propose a semiparametric MoE model in which each expert error distribution is represented as a nonparametric Gaussian scale mixture estimated via nonparametric maximum likelihood, relaxing parametric assumptions within the Gaussian scale-mixture class while preserving the interpretability and structure of the MoE framework. The resulting model adapts to complex error structures, improves robustness under contamination and heavy tails, and remains competitive under well-specified Gaussian settings, providing a practical and theoretically grounded alternative to parametric MoE formulations.
Optimal minimization of an unknown function in a nonparametric multivariate regression model thanks to a dimension reduction approach
oai:arXiv.org:2608.04566v1
arXiv:2608.04566v1 Announce Type: new
Abstract: In this paper, we propose a novel approach for estimating the minimum of a smooth function and its location from observations corresponding to a multivariate regression function depending a priori on d variables but actually only on r < d active variables and corrupted by some additional noise. Our method consists of two steps: The rst one is a variable selection approach which is used for identifying the r active variables on which f depends and the second one consists in estimating the minimum of the function and its location. The estimation of the minimizers is obtained by using a projected gradient descent where the gradient is estimated using a local polynomial approximation of the regression function limited to its active variables obtained in the rst step. The estimation of the minimum is obtained by evaluating the estimator of the regression function using a local polynomial approach at the estimator of one of the minimizers previously obtained. We establish non asymptotic upper bounds for the quadratic risk of the estimators of the minimizers and of the minimum and prove that they reach the optimal rate that could be expected as if the active variables were known beforehand up to a factor smaller than a power of a logarithmic term.
Automatic Statistical Test for Rationally Expressible Algorithms by Selective Inference, with Applications to Feature Selection
oai:arXiv.org:2608.04667v1
arXiv:2608.04667v1 Announce Type: new
Abstract: Selective inference (SI) provides statistically valid $p$-values for hypotheses selected by applying an algorithm to the data, correcting for the bias that arises when the same data are used both to select and to test a hypothesis. Developing an SI procedure for a new algorithm, however, has required an expert to derive, and then implement, the selection event, i.e., the conditions under which the hypothesis is selected. Repeating this specialized effort for every new algorithm is why exact SI has so far been available for only a narrow class. We propose AutoSI, a framework that removes this barrier in two ways. First, AutoSI constructs the selection event automatically from the algorithm's individual operations, so the user only writes the algorithm as ordinary NumPy-like code and derives nothing by hand. Second, AutoSI broadens the class of selection events SI can handle: existing exact methods are limited to selection events characterized by linear or quadratic inequalities in the data, whereas AutoSI covers any algorithm expressible through rational functions of the data (ratios of polynomials). We prove that the $p$-values computed by AutoSI are exactly valid in finite samples. We demonstrate AutoSI on three feature-selection methods, each written in a few dozen lines of code. One of these methods, the lasso with its tuning parameter selected by cross-validated $R^2$, cannot be handled within existing exact SI frameworks and is made possible by AutoSI. Experiments on synthetic and real datasets show that the resulting $p$-values control the type I error rate (i.e., the false positive rate) at the nominal level while retaining high power.
Consistent community recovery in stochastic block Ornstein-Uhlenbeck processes
oai:arXiv.org:2608.04700v1
arXiv:2608.04700v1 Announce Type: new
Abstract: We propose the stochastic block Ornstein-Uhlenbeck (SBOU) process, a continuous-time multivariate model in which the drift matrix encodes a latent group structure among its components. Our main contribution is a community-detection algorithm whose misclassification proportion converges to zero in a regime combining infill, long-span, and high-dimensional asymptotics. To our knowledge, this is the first consistency result of this kind for latent group recovery in a discretely observed continuous-time multivariate model. As a key intermediate result, we establish consistency of the discretely observed maximum likelihood estimator of the drift matrix in the same regime, thereby extending the high-dimensional L\'evy-driven Ornstein-Uhlenbeck literature. For practical implementation, we develop a feasible model-selection procedure for estimating the support of the drift matrix, which enables data-driven selection of the number of latent groups. The SBOU framework can be viewed as a continuous-time generalisation of the discrete-time stochastic-block VAR model, allowing for both positive and negative dynamic interactions between groups as opposed to only positive. We illustrate the methodology on the RE-Europe wind-capacity dataset and recover a country-level grouping consistent with the geographic benchmark.
Debiasing the Lasso under Weaker Tail Assumptions
oai:arXiv.org:2608.04800v1
arXiv:2608.04800v1 Announce Type: new
Abstract: We consider the problem of high-dimensional inference with the lasso estimator. Different methods including 'double selection' techniques and multiple versions of the 'debiased lasso' have been proposed for this task with noticeable success. However, most guarantees assume strong hypotheses on the underlying data process and the errors in the linear regression model, such as subgaussian designs and independence between errors and the data itself. We show that 'standardizing' one's dataset -- a natural procedure in practical penalized regression -- leads to same results under much weaker hypotheses, paying only a small price for not assuming light tails. The key technical point allowed by this step is exploiting the concentration properties of self-normalized processes. Importantly, we prove our results for two different methods closely related to the 'debiased lasso'. The second method performs valid inference even for a misspecified linear model, under mild sparsity conditions similar to the 'double selection' literature.
Constructing Large Orthogonal Minimally Aliased Response Surface Designs Through Enumeration and Combination of Weighing Designs
oai:arXiv.org:2608.04814v1
arXiv:2608.04814v1 Announce Type: new
Abstract: Advances in automation and high-throughput experimentation have enabled larger and more complex studies involving many factors and tests, creating a growing demand for computationally effective design construction methods. Efficient experimental design remains a key challenge in this context, creating a need for frameworks that can generate large experiments while preserving orthogonality and minimal aliasing. Unlike existing approaches which struggle with scalability, this work introduces an algorithmic framework for constructing large Orthogonal Minimally Aliased Response Surface (OMARS) designs by enumerating and combining weighing designs, three-level matrices with orthogonal columns and a fixed number of non-zero entries per column. Complete enumerations of weighing designs are achieved for designs with up to 24 tests, covering multiple numbers of factors and weights corresponding to two or three zeros per factor. In addition, a validated partial enumeration procedure and a combination method extend the catalog to substantially larger designs. The combination method enables the construction of OMARS designs for any test size that is a multiple of selected base sizes. This paper thus provides the methodology for generating large catalogs of high-quality OMARS designs, well-suited for high-dimensional screening and response-surface modelling in complex industrial and scientific experiments.
Group-regularized matrix factorization for fast and reliable module discovery in pan-omics pan-cancer studies
oai:arXiv.org:2608.04826v1
arXiv:2608.04826v1 Announce Type: new
Abstract: In pan-omics pan-cancer studies, it is critical to identify latent sources of variation that are shared across particular subsets. This task often requires bidimensionally linked data matrices to be decomposed into a sum of block-sparse, low-rank modules. Existing approaches often rely on pre-specified module numbers, ranks, or post-hoc thresholding and can be sensitive to model specification when the underlying sharing structure is complex. To address these issues, we propose GL-BIDIFAC+, a group-regularized matrix factorization framework for discovering partially shared modules. It requires only an upper bound on the latent dimension and encourages module selection through group regularization with theoretically-motivated tuning parameter selection and local support recovery analysis, providing both scalability and principled guidance for module discovery. It also admits a probabilistic interpretation that enables model-based imputation of missing data. Simulation studies demonstrate accurate module recovery and favorable computational performance relative to existing approaches. We further apply GL-BIDIFAC+ to analyze the Cancer Genome Atlas data, where well-established molecular structure provides interpretable biological references. Our analysis distinguishes broad pan-cancer variation, cancer-specific subtype structure, and variation shared across cancers with related tissue origins or histologic features.
Intrinsic-Hybrid Latent Diffusion Models for Generative Modeling on Unknown Manifolds
oai:arXiv.org:2608.04827v1
arXiv:2608.04827v1 Announce Type: new
Abstract: We introduce the Intrinsic Hybrid Latent Diffusion Model (ILDM), a generative framework that integrates probabilistic dimensionality reduction with geometry-aware diffusion on unknown manifolds. While diffusion models (DMs) have achieved state-of-the-art results in high-dimensional data synthesis, they rely on large training datasets and ignore intrinsic geometric structure. Latent diffusion models (LDMs) address the high dimensionality by learning a latent space, but they typically impose a Euclidean structure, failing to capture the underlying manifold geometry, especially problematic in data-sparse regimes. ILDM addresses these limitations by interpreting the latent space as a chart of an unknown Riemannian manifold, with geometry and uncertainty quantified through a probabilistic decoder. The forward process is a hybrid diffusion that switches between Riemannian and Euclidean dynamics based on local uncertainty, where the Riemannian component is governed by a probabilistic metric tensor derived from the decoder. To learn the generative dynamics, we introduce an approximate denoising score matching method tailored to the hybrid diffusion setting, enabling a backward process defined by hybrid Langevin dynamics. Experiments on COIL-100, MNIST, and cardiac MRI datasets demonstrate that ILDM significantly improves generation quality, achieving lower FID and LPIPS scores compared to standard diffusion and latent diffusion models.
Exchangeable Testing Against an Unknown Benchmark
oai:arXiv.org:2608.04838v1
arXiv:2608.04838v1 Announce Type: new
Abstract: We generate infinite binary exchangeable sequences by sequential comparison of data points against a latent benchmark. Assuming a prior distribution of the benchmark rank \(R_0\) within an unobserved group, we set up the Bayesian machinery that determines the posterior distribution of the running rank \(R_n\) in purely combinatorial terms. This yields an explicitly computable predictive probability of winning against the benchmark. The normalised running rank converges to a latent strength variable \(X\) with polynomial density, possibly Beta-tilted.
Some min-max tournaments lead to particularly simple multiplicative formulae for predictive probabilities related to priors that generalise the Topp--Leone distribution; for that class we analyse the asymptotics of the associated fixed-\(n\) up-down Markov chains. The limiting diffusion has the classical Wright--Fisher variance but a nonlinear drift expressed explicitly via the prior density of the benchmark.
Mixtures of Beta densities are classical objects in the theory of exchangeable sequences. The contribution of the present work is the combinatorial rank-based updating mechanism and the resulting explicit predictive laws for sequential testing against an unknown benchmark.
Nonparametric Goodness-of-fit Testing under Covariate Shift
oai:arXiv.org:2608.04860v1
arXiv:2608.04860v1 Announce Type: new
Abstract: This paper develops procedures for nonparametric goodness-of-fit testing under covariate shift, where labelled data are drawn from a source population but goodness-of-fit is evaluated for a target population. The distribution mismatch is quantified by either a bounded moment condition or a sub-exponential tail condition on the target-to-source density ratio. Our method combines truncated importance-weighting kernel ridge regression with a multiplier bootstrap to construct confidence sets for the regression function. The truncation stabilizes the importance- weighting kernel ridge regression as well as the bootstrap calibration, making our approach applicable even when the density ratio has heavy tails. We prove nonasymptotic validity and sharpness of the resulting confidence sets under suitable operator compatibility conditions, and establish explicit error rates for coverage probability under specific conditions on the target- to-source density ratio and on the spectral decay of the kernel integral operator. Numerical experiments corroborate our theoretical findings.
A Design-Based Minimax Theory for Network Experiments
oai:arXiv.org:2608.04909v1
arXiv:2608.04909v1 Announce Type: new
Abstract: Network experiments are used throughout the social and medical sciences to investigate causal effects under the presence of interference. While a large body of work has developed improved statistical procedures, the fundamental limits of statistical estimation in these settings is less well understood. In this paper, we develop and investigate a design-based theory of minimax risk for network experiments under an arbitrary neighborhood interference model. Our notion of minimax risk describes the optimal precision among all statistical procedures for investigating a particular causal effect on the observed interference network. We show that the minimax risk is a function of the corresponding conflict graph, which captures inherent unobservability of estimand-relevant potential outcomes given the observed interference network. Our main contribution is a series of upper and lower bounds on the minimax rate in terms of local and global connectivity properties of the conflict graph. To illustrate their utility, we apply these general results to obtain minimax analyses for two commonly studied effects: the direct treatment effect and global average treatment effect.
Inverse probability weighting for auxiliary variable dependent sampling in observational studies of Long COVID
oai:arXiv.org:2608.04918v1
arXiv:2608.04918v1 Announce Type: new
Abstract: Selective testing based on values of auxiliary variables is an increasingly popular design strategy in observational studies and is ubiquitous in electronic health record data. Ignoring this underlying sampling mechanism can lead to biased estimation and erroneous scientific conclusions. Yet, rigorous analytic methods for accounting for two-phase sampling designs in observational settings remain under-utilized. Motivated by the Researching COVID to Enhance Recovery (RECOVER) Adult and Pediatric observational cohort studies, we describe common pitfalls and an approach for analysis of data collected via auxiliary variable dependent sampling.
Statistical Considerations in Long COVID Research
oai:arXiv.org:2608.04919v1
arXiv:2608.04919v1 Announce Type: new
Abstract: Long COVID is a condition characterized by ongoing or relapsing symptoms attributable to SARS-CoV-2 infection that are present three or more months after infection. It represents a major clinical and public health concern as an estimated 5-10\% of individuals with a history of SARS-CoV-2 infection present with long term sequelae that range from mild to debilitating with profound impacts on quality of life. Clinical research studies of Long COVID have emerged rapidly over the past few years, and with them we are seeing several new data analytic challenges. In this manuscript, we highlight statistical challenges arising from the defining features of LC and associated study design strategies. This work is motivated by the Researching COVID to Enhance Recovery (RECOVER) Adult and Pediatric observational meta-cohort studies.
Parameter identification for predator-prey system with sparse data
oai:arXiv.org:2608.04959v1
arXiv:2608.04959v1 Announce Type: new
Abstract: Parameter identification from observations of dynamical systems is a fundamental problem in population biology. Mechanistic models of ecological systems rely on optimization methods that require accurate initial guesses to guarantee convergence. In ecological applications, datasets contain observation noise and are collected at sparse time points. This sparsity creates irregular likelihoods that cause standard optimization methods to struggle, while the ordinary differential equation solvers can become stiff or unstable in certain regions of the parameter space. These instabilities cause long running times or runtime errors. Here we present a computational framework for parameter identification that addresses these numerical instabilities by employing Natural Gradient Ascent, and we apply it to the classical Lotka-Volterra predator-prey model. We exploit the non-dimensionalization of the ordinary differential equations to treat scaling factors as nuisance parameters, reducing the dimensionality of the optimization problem. To prevent the solver step from becoming small, we implement an adaptive solver that switches between two independent second-order equations derived from the two components of the model. This approach allows Natural Gradient Ascent to converge in fewer iterations and with more stability than standard gradient ascent or BFGS methods. This framework provides a reliable method for parameter estimation in ecology when data is limited. The method can be generalized to other dynamical systems as long as the different components of the system do not become numerically problematic at the same time.
HaploPerturb: Low-rank copula construction of haplotype perturbations improves sequence-to-function analysis of Alzheimer's disease loci
oai:arXiv.org:2608.05002v1
arXiv:2608.05002v1 Announce Type: new
Abstract: Sequence-to-function models predict molecular phenotypes from complete sequence windows. At genome-wide association study loci, however, the prevailing design perturbs only the lead variant on the reference genome, even though the lead is often correlated with nearby variants through linkage disequilibrium. This single-variant perturbation implicitly fixes all linked alleles at their reference-genome states and may therefore create an uncommon or unobserved population haplotype. We study this input-construction problem at 38 Alzheimer's disease loci. We introduce HaploPerturb, which uses phased ROS/MAP genotypes or the European 1000 Genomes panel to fit a fixed-margin latent Gaussian factor model and rank partner configurations conditional on each lead allele. The leading public-panel construction agrees with the donor-panel construction at all loci under a strict linkage-disequilibrium threshold and 36 of 37 loci under a broader threshold after restricting to shared partners. Known-truth simulations show exact recovery of the dominant configuration under strong linkage disequilibrium and expose persistent residual correlation under a misspecified one-factor model. In an AlphaGenome benchmark against cell-type-specific ROS/MAP eQTLs, broader-set public-panel haplotypes yield microglial enrichment of 2.07 (95\% whole-locus bootstrap percentile interval 1.48--3.60), compared with 1.43 (0.69--2.29) for a lead-only edit. Empirical-mode and LD-sign backgrounds yield 2.20 (1.60--3.64), with no detectable advantage or loss relative to the HaploPerturb top configuration. Thus population-informed sequence construction matters in this application, while the choice among reasonable leading haplotype rules is less consequential than the choice between a haplotype and a lead-only reference background.
Analyzing the daily flows: Exploring shared micro-mobility factors in Venice
oai:arXiv.org:2608.05065v1
arXiv:2608.05065v1 Announce Type: new
Abstract: Shared micro-mobility has emerged as a key component of a sustainable urban transportation system, however, limited research exists on how environmental factors influence the mobility demand between specific origin-destination (OD) locations. This work extends research on the demand-side perspective to explore how temporal and environmental conditions shape daily shared micro-mobility flows in Venice. The study analyses repeated variation across 158,401 OD-day observations for two years in 50 spatial zones. Daily temperature, rainfall, and PM10 concentrations are linked to each OD-day observation while accounting for vehicle-pass composition and temporal patterns. Here, the unit of analysis is the connection between OD pairs. A generalised additive mixed model (GAMM) is used to represent the non-linearity of environmental relationships across seasons, providing a flexible framework for understanding how climatic conditions influence sustainable mobility behaviour. The results show a significant nonlinear association between temperature and mobility demand across seasons. High rainfall is associated with reduced demand, with larger reductions under moderate and heavy rainfall than on dry days. The relationship between PM10 and shared mobility use was season-dependent, creating an avoidance-versus-adoption mechanism rather than a monotonic association. After adjustment for environmental and temporal factors, a recurring increase in demand within the Lido Islands during August and September remained evident, highlighting a location-specific mobility pattern across two years. The study highlights the importance of environmental sensitivity in shared micro-mobility research. This work illustrates that the adoption of shared bikes and electric bikes depends not only on service availability but also on usage patterns, which are affected by external conditions.
Nonparametric Estimation under General Nonlinear ODE Constraints: A Comparison with Parametric ODE-Fitting Methods
oai:arXiv.org:2608.05081v1
arXiv:2608.05081v1 Announce Type: new
Abstract: Many physical, biological, and epidemiological processes are governed by ordinary differential equations (ODEs) that are nonlinear in the state variable, including logistic population growth, chemical reaction kinetics, and epidemiological compartment models. We develop a differential equation-constrained local polynomial regression (DE-constrained LPR) framework for the general first-order ODE constraint g'(x) = F(x, g(x)), where F may be any Lipschitz continuous function, extending prior work restricted to exponential and linear ODE structures. Because F is generally nonlinear in g, the Taylor coefficients of the DE1-k estimator cannot be written in closed form; instead they are obtained by successive symbolic differentiation of F, and the estimator is computed by nonlinear least squares, requiring only a single local parameter at each evaluation point regardless of polynomial degree k. We derive the asymptotic conditional bias and variance of the DE1-k estimator, propose an AIMSE-optimal bandwidth that exploits the ODE structure to avoid direct estimation of high-order derivatives, and evaluate the method in a simulation study based on logistic growth, benchmarking against the parameter cascading method of Ramsay et al. (2007) (PCODE) and classical local linear regression. The DE-constrained estimator consistently outperforms local linear regression and is competitive with PCODE even though it estimates no structural parameter of the ODE; a sensitivity analysis across growth rates shows DE-constrained estimation becomes more accurate and more robust than PCODE as the curve steepens and PCODE's parameter estimation grows less stable. These results position DE-constrained LPR as a practical nonparametric alternative to parametric ODE-fitting methods when structural parameters are difficult to identify reliably.
Stable Density Ridges: Consistency and Convergence of Subspace Constrained Mean Shift
oai:arXiv.org:2608.05112v1
arXiv:2608.05112v1 Announce Type: new
Abstract: The Subspace Constrained Mean Shift (SCMS) algorithm is a popular nonparametric method for extracting density ridges, which serve as a low-dimensional representation of high-dimensional data. It is a widely held belief in the literature that SCMS trajectories converge to the classical density ridge, which we call the "static ridge", defined via the density gradient and the eigenvalues and eigenvectors of the density's Hessian. In this paper, we demonstrate that this assumption does not hold in general, as the static definition fails to account for the rotation of the trailing eigenspace along the continuous flow of the algorithm's underlying vector field. To resolve this, we propose a paradigm shift by introducing the "stable ridge", a novel geometric structure defined through the lens of dynamical systems and the Jacobian of the projected density gradient. We prove that this stable ridge is the true theoretical target of the SCMS algorithm. Building upon this foundation, we develop a generalized SCMS framework utilizing a constant step size, establishing its uniform R-linear convergence and topological surjectivity onto the stable ridge. We further derive the rates of convergence for estimating the stable ridge in terms of the Hausdorff distance. Finally, we expose that the original SCMS algorithm suffers from polynomial-time computational complexity, which is caused by implicitly coupling the step size to the smoothing bandwidth via the Mean Shift operator, and demonstrate how our generalized framework provides a statistically consistent and more efficient solution.
Micro-randomized Trials with Categorical Treatments and Binary Proximal Outcome: Causal Effect Estimation and Sample Size Calculation
oai:arXiv.org:2608.05135v1
arXiv:2608.05135v1 Announce Type: new
Abstract: Micro-randomized trials (MRTs) provide a framework for evaluating the marginal and moderated effects of mobile health (mHealth) interventions. In many applications, treatments take the form of categorical variables with multiple levels, such as different message contents or delivery strategies. Many scientifically meaningful longitudinal outcomes in mHealth studies are binary, such as whether a participant opens an app, engages with content, or completes a target behavior following a decision point at which treatment is randomized. This paper focuses on MRTs with categorical treatments and binary proximal outcomes. We define the causal excursion effect, propose an estimator called EMEE-catA, and derive a sample size formula for comparing categorical treatment levels that controls the type I error rate and guarantees power under working assumptions. We conduct extensive simulation studies to evaluate the operating characteristics of the proposed sample size formula, including robustness to violations of these assumptions. We further provide practical guidance for implementing the proposed approach to ensure adequate power in real-world MRTs. The methods are illustrated using data from the Drink Less MRT.
Informational Frustration in Neural Manifolds: Shannon Bottlenecks and the Limits of Learnability
oai:arXiv.org:2606.30512v1
arXiv:2606.30512v1 Announce Type: cross
Abstract: Why overparameterised deep networks generalise so remarkably well remains one of the most stubborn open questions in machine learning theory. Classical frameworks like VC dimension and Rademacher complexity predict catastrophic overfitting in modern models, leaving a massive theoretical gap between theory and reality. In this paper, we bridge this divide by introducing a unified framework that links information theory, topology, and statistical mechanics to map the hard limits of deep learning. Central to our approach is the Entropic Learnability Horizon (ELH): a fundamental law stating that a network can only truly learn a target function if the Shannon entropy of the data manifold outpaces the topological entropy of the function's decision boundary, balanced by the von Neumann entropy of the network's weight space. We establish the Shannon-Topological Bottleneck Theorem, proving that when a target boundary's geometric complexity exceeds this informational horizon, the system undergoes a sudden entropic phase transition. It falls into a state of Informational Frustration - a glassy, rigid memorization phase where generalization becomes thermodynamically impossible. Using this lens, we show that the enigmatic phenomenon of "grokking" is actually an Entropic Release, where weights abruptly reorganise to unlock the bottleneck. Finally, we translate this theory into practice with Entropic Gradient Descent (EGD), an optimization algorithm that dynamically manages weight entropy to keep learning on track. Ultimately, this work repositions entropy not just as a tool for tracking uncertainty but as the fundamental physical currency that dictates whether a machine can learn.
Statistical Mechanics of Learning on Product Wasserstein Manifolds
oai:arXiv.org:2608.01434v1
arXiv:2608.01434v1 Announce Type: cross
Abstract: Normally the statistical mechanics of learning treats constraints on weight distributions as restrictions that shrink the space of possible solutions. Therefore, it reduces model capacity. In this paper we would like to take a contrary approach, which, however, is based on the earlier work on distribution-constrained perceptrons. Rather than treating a prescribed weight distribution as a mere restriction, we propose that it defines the intrinsic geometry upon which learning naturally unfolds. We formulate both deep neural networks and variational quantum circuits as gradient flows on a product of Wasserstein manifolds -- one classical Wasserstein space for each layer and one quantum Wasserstein space for the circuit parameters. Within this geometry, the capacity reduction, which was previously associated with distributional constraints, appears as the metric structure of the constraint manifold itself. We develop a hierarchical mean-field description for deep networks, extend the framework to the quantum setting using the quantum Wasserstein distance of order 1, and introduce two such practical algorithms, Hierarchical DisCo-SGD and Quantum DisCo, that follow approximate geodesics on the manifold of the product itself. Experiments on teacher-student problems, standard image classification tasks, and small variational quantum classifiers show that respecting these distributional geometries improves generalization, stabilizes training, and reduces the severity of barren plateaus compared with unconstrained and purely norm-based baselines. This approach firstly reframes structural constraints as geometric priors and suggests a route for incorporating biological, spectral, or hardware-derived distributional information into both learning systems, viz., classical and quantum learning.
AI-driven Multimodal Representation Learning for Latent Mediation Structure Discovery of Socioeconomic Disadvantage, Psychosocial Factors, and Cardiometabolic Multimorbidity: Insights from the All of Us Research Program
oai:arXiv.org:2608.04016v1
arXiv:2608.04016v1 Announce Type: cross
Abstract: Social disadvantage is associated with multimorbidity, but the pathways linking social conditions to disease burden remain poorly understood. We developed an AI-driven multimodal mediation framework that integrates socioeconomic, psychosocial, clinical, laboratory, behavioral, and genomic data from the All of Us Research Program. Modality-specific variational autoencoders were used to derive latent representations of each data domain, and mediation analyses were subsequently performed in latent space to evaluate indirect associations between socioeconomic disadvantage, psychosocial factors, and multimorbidity. The final analytic cohort included 20,804 participants with complete multimodal data. Across 800 exposure--mediator--outcome combinations, mediation signals were concentrated within a small number of latent dimensions. The strongest indirect association linked a socioeconomic disadvantage dimension, a psychosocial vulnerability dimension, and a cardiometabolic multimorbidity dimension (NIE = 0.002517). The psychosocial dimension was characterized by poorer mental health, greater loneliness, lower social well-being, and lower health literacy, whereas the outcome dimension was associated with hypertension, diabetes, hyperlipidemia, obesity, chronic kidney disease, and heart disease. Bootstrap analyses supported the stability of the leading pathway. These findings suggest that psychosocial vulnerability was strongly represented in the dominant latent pathway linking socioeconomic disadvantage and cardiometabolic multimorbidity. More broadly, the proposed framework illustrates how AI-based representation learning can be used to investigate complex relationships across high-dimensional multimodal health data.
Forced Displacement of People Experiencing Homelessness: Housing and Movement Outcomes after Encampment Clearances
oai:arXiv.org:2608.04076v1
arXiv:2608.04076v1 Announce Type: cross
Abstract: The 2024 Grants Pass decision newly emboldens US cities to manage unsheltered homelessness through forced displacement. Although literature demonstrates the harmful health and material impacts of this tactic, the longer-term results for housing outcomes and migration patterns remain less clear. In response, this study leverages longitudinal street outreach data to investigate where people move following encampment clearances. We specifically employ relational event models to predict the likelihood of various outcomes post-removal, such as relocating tracts, entering shelter, or obtaining housing. Results suggest that displaced residents do not travel far, yet clearances may still reduce visible homelessness by decreasing the size of camp communities. Furthermore, people appear unlikely to move indoors and instead face high risks of losing contact with service providers. These trends hold regardless of individual demographics, although people with mental health conditions demonstrate stronger attachments to their original sites. Evidence additionally indicates that neighborhood conditions could influence these migration behaviors. Such findings corroborate broader literature on place attachments, residential mobility, and invisibilization of poverty. This paper ultimately addresses the urgent need for a deeper understanding of how forced displacement impacts homelessness and pathways to housing.
Correlation Matrices in High Dimensions: The Elliptope as a Sample-Correlation Ensemble
oai:arXiv.org:2608.04162v1
arXiv:2608.04162v1 Announce Type: cross
Abstract: The set of $n\times n$ correlation matrices, known as the elliptope, has volume decaying at the super-exponential rate $\exp\{-\tfrac14 n^2\log n\}$. We characterize where this vanishing volume concentrates. A uniform draw is entrywise close to the identity yet globally far from it and nearly singular: its maximum absolute correlation is of order $\sqrt{\log n/n}$, its Frobenius distance is asymptotic to $\sqrt n$, its empirical spectral distribution converges to the Marchenko-Pastur law with ratio one, and its smallest eigenvalue has the exact $\operatorname{Beta}(1,d)$ distribution, where $d=n(n-1)/2$, and is therefore of order $n^{-2}$. More generally, distinct off-diagonal entries are exactly pairwise independent under every $\operatorname{LKJ}(\eta)$ law. For the uniform law, this yields a Chen-Stein proof of the extreme-correlation point-process limit and an $O(n^{-1})$ total-variation bound for finite-dimensional exceedance counts relative to Poisson laws with their exact finite-$n$ means. We also identify two distinct scales: $\eta_n\asymp n$ alters the limiting spectrum, whereas $\eta_n\asymp n^2$ is needed to keep the Frobenius distance bounded. Finally, for a bounded, centered i.i.d. off-diagonal specification, projection to the nearest correlation matrix incurs a squared repair cost asymptotically at least one-half of the squared Frobenius norm of its off-diagonal part.
Sample Complexity of Multicalibration for Multilevel Properties
oai:arXiv.org:2608.04288v1
arXiv:2608.04288v1 Announce Type: cross
Abstract: Calibration requires a predictor to be unbiased after conditioning on its own predictions. Multicalibration asks for this guarantee simultaneously across a collection of groups. Many prediction tasks ask for several related features of the same conditional outcome distribution: variance is defined relative to the mean, skewness relative to both mean and variance, and conditional value at risk relative to a quantile. We study multicalibration for a sequence of $k$ properties in which each property is identifiable once the preceding properties are fixed. This framework includes Bayes pairs but does not require the properties to arise from a single loss.
For every fixed $k\ge2$, we establish matching upper and lower sample-complexity bounds up to logarithmic factors under regularity conditions. Even with only polylogarithmically many binary groups, achieving multicalibration error $\varepsilon$ requires $\widetilde{\Omega}(\varepsilon^{-(k+2)})$ samples. Conversely, for any finite group family $\mathcal G$, we give a randomized learner using $O(\varepsilon^{-(k+2)}+\varepsilon^{-2}\log|\mathcal G|)$ samples. Thus the sample complexity is $\widetilde{\Theta}(\varepsilon^{-(k+2)})$ for polynomial-size group families. We instantiate the theory for three canonical examples.
ArborEnum: Decision Tree Rashomon Sets over Continuous Features
oai:arXiv.org:2608.04310v1
arXiv:2608.04310v1 Announce Type: cross
Abstract: The Rashomon effect describes the phenomenon that many models can achieve nearly equivalent performance on the same learning task, with significant ramifications for robustness, feature importance, and customizability. These use cases motivate the computation of Rashomon sets: the set of all models whose regularized loss is near-optimal. Decision trees are one of the few model classes for which Rashomon sets can be fully enumerated, but this computation has always been conditional on a binarization of the original data, either restricting which splits each tree is allowed to make or substantially increasing the complexity of an already difficult combinatorial problem. We introduce the first algorithm that exactly enumerates decision-tree Rashomon sets while exploiting the ordered structure of continuous features. We further develop a relaxation for approximate enumeration and an anytime algorithm that progressively refines the set of candidate thresholds, producing increasingly detailed approximations that converge to the continuous-feature Rashomon set. Experiments show that coarse binarization can miss many trees, important features, and predictive multiplicity; our algorithms achieve orders-of-magnitude speedups over existing enumeration methods, with approximations providing further speedups while maintaining near-perfect recall.
Achieving First-Order Statistical Improvements in Data-Driven Optimization: From No-Free-Lunch to Amplified Decision Perturbation
oai:arXiv.org:2608.04312v1
arXiv:2608.04312v1 Announce Type: cross
Abstract: Recent proliferation of data-optimization integration has led to a range of methods that aim to improve the statistical performance of data-driven optimization decisions. However, while many of these methods are motivated intuitively from a robustness or regularization perspective, their resulting statistical benefits are often unclear and, even if available, are established on a case-by-case basis. We provide a systematic dissection of data-driven optimization formulations using the view of "directionally perturbed" empirical optimization (EO). Specifically, this umbrella of formulations, which we call "EO+", covers many existing data-driven optimization methods, including regularization, distributionally robust optimization, transfer learning, and analogous methods for contextual optimization. On the one hand, we argue that without additional, correctly specified, side information, any EO+ method can result in at most second-order improvements. This provides a negative conclusion, namely ``no free lunch is possible", on the statistical power of EO+. On the other hand, we show that when leveraging side information that is geometrically effective, achieving first-order improvements is possible by choosing hyperparameters that are significantly larger than what is typically suggested in the literature. Moreover, we construct a principled methodology based on excess risk estimation, via either system knowledge or bootstrap resampling, to maximize the first-order gain. We demonstrate how this gain connects to the control-variate principle, a variance reduction technique in the Monte Carlo simulation literature, which helps explain why geometrically effective side information is necessary.
Equitable System-Prompt Selection via Constrained Mixed-Strategy GroupDRO
oai:arXiv.org:2608.04339v1
arXiv:2608.04339v1 Announce Type: cross
Abstract: Large language models are increasingly used for information seeking, yet semantically equivalent questions phrased in different ways can receive answers of considerably different quality. System prompts are widely employed to steer response behavior, but they are typically optimized for average-case quality, so some question phrasings may still receive incomplete or low-quality answers. To address this, we formulate a constrained mixed-strategy GroupDRO framework for system-prompt selection. Instead of optimizing the system-prompt text, the framework assigns weights to system prompts in an existing pool to minimize the worst-case information-quality loss across evaluation metrics and groups, while constraining the mean loss to stay close to that of average-based selection. Because pool generation and selection are decoupled, the method applies to any system-prompt pool and can leverage an ensemble of complementary system prompts rather than a single one. Across five LLMs on two bilingual medical and consumer-finance benchmarks, the constrained method reduces the Overall Mean, Worst 25% Mean, and Worst by 13.1%, 13.2%, and 13.7% on average relative to no mitigation while keeping overall quality close to Average selection. Its multi-prompt weights reveal complementarity across metric-group pairs. Code and data are available at https://github.com/Rainxu09/equitable-system-prompt-selection.
iStructTab: Structured Feature Sequencing for Multimodal Learning of Image and Tabular Data
oai:arXiv.org:2608.04348v1
arXiv:2608.04348v1 Announce Type: cross
Abstract: Multimodal learning of images and tabular data is often impaired by ineffective representations, resulting in redundancy, dispersion, and generalization problems. To tackle this challenge, we introduce Graph-Enhanced Descriptor Sequencing (GEDS), a structured feature sequencing algorithm grounded in principles from the Column Permutation Problem (CPP). GEDS refines statistical descriptors of the features through similarity graph-based computations, systematically determining an effective feature sequencing. We incorporate GEDS within an order-aware efficient transformer framework, utilizing order-aware memory tokens that explicitly adhere to the derived feature sequencing via a dedicated loss function. Experimental results across multimodal benchmarks demonstrate that iStructTab effectively minimizes feature dispersion, improving predictive performance and robustness, and highlighting the significance of structured feature sequencing in multimodal learning.
Non-asymptotic implicit bias of logistic regression at early-stage gradient descent dynamics
oai:arXiv.org:2608.04382v1
arXiv:2608.04382v1 Announce Type: cross
Abstract: Gradient descent has been of particular interest in modern machine learning beyond sole focus on optimization. Implicit bias emerging from optimization, though not being encoded by the learning objective, often prevents from overfitting to spurious patterns. A typical instance is the max-margin implicit bias of a linear classifier, widely established for exponentially tailed loss functions. Even after having a given dataset separated, the parameter vector continues to evolve towards the max-margin direction asymptotically along the gradient descent dynamics. This phenomenon corroborates a frequent empirical observation of "train longer, generalize better." However, the max-margin convergence is an asymptotic phenomenon, and what is worse, this asymptotic convergence rate is significantly slower than pure convex optimization. Even so, the parameter vector along gradient descent dynamics commonly correlates with the max-margin direction positively (though not exactly) within considerably fewer iterations than the asymptotic rate. By shedding another light on this classical problem, this work aims to understand the mechanism of this early-stage alignment phenomenon. Our theoretical results demonstrate that the parameter vector weakly aligns with the max-margin direction within $O(\exp(\exp(-\delta)))$ iterations, where $\delta>0$ is the permissible alignment error, which is shown to be tight. By tracking the radial and tangential flows, our proof operates on the alignment dynamics directly with dataset geometry and gets rid of the asymptotic expansion, which is a key insight to establishing faster weak alignment.
Incremental Aggregation on the Grassmannian for Asynchronous Eigenspace Computation
oai:arXiv.org:2608.04406v1
arXiv:2608.04406v1 Announce Type: cross
Abstract: We study asynchronous optimization for finite-sum eigenspace computation in heterogeneous distributed systems. The theoretical foundations for asynchronous eigenspace computation remain scarce, with existing approaches offering limited coverage of dynamics directly on the Grassmannian under stale information. In this paper, we propose a Grassmannian incremental aggregation method that refreshes only arriving components and reuses cached gradients, retaining low per-update cost without global synchronization. The method employs an extrinsic polar update that preserves the intrinsic subspace geometry without requiring parallel transport of stale tangent vectors. Our analysis establishes a tight angle-dependent gradient-dominance characterization of the objective and a basin-invariance property for stale aggregated updates. These yield two-phase linear convergence, comprising an explicit broad-basin regime and a sharper local regime, with constants controlled by component spectral spreads. Experiments on serial and distributed PCA demonstrate improved sample efficiency and wall-clock convergence over representative baselines.
An adaptive split-combine Gaussian mixture filter for nonlinear and multimodal state estimation
oai:arXiv.org:2608.04430v1
arXiv:2608.04430v1 Announce Type: cross
Abstract: Filtering combines model predictions with measurements to estimate the probability density function (PDF) of a system state over time. The PDF often becomes highly asymmetric and even multimodal in nonlinear systems with oscillatory or chaotic dynamics. Such non-Gaussian features violate the single-Gaussian assumption underlying Kalman-type filters. To address this problem, Gaussian mixture filtering has been proposed. However, accurately propagating mixture components and adaptively adjusting their number and weights over time remain open challenges. Here, we develop an adaptive split-combine Gaussian mixture filter (AMF) that estimates the time evolution of asymmetric and multimodal PDFs by adaptively splitting and combining Gaussian particles without auxiliary online numerical optimization. Notably, the proposed splitting method guarantees a reduction in variance along a target level-set-point direction of a Gaussian particle. This enables accurate and efficient propagation of particles. We show that AMF consistently outperforms various baseline filters across diverse benchmarks, including single and coupled slow-fast Van der Pol oscillators and the Lorenz attractor. We also propose a parallel implementation of AMF, allowing high-fidelity PDF estimation with practical computational cost.
The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
oai:arXiv.org:2608.04432v1
arXiv:2608.04432v1 Announce Type: cross
Abstract: On two-sided content platforms, symmetric two-sided isolation (assigning matched fractions of creators and viewers to isolated treatment and control submarkets) is widely used for creator-side and cold-start experiments because it removes cross-arm marketplace interference. Isolation, however, thins each viewer's candidate catalog, and intuition suggests the resulting engagement cost should fade as the platform grows: a small fraction of a vast catalog is still vast. We show that, in an order-statistics model of engagement, whether this intuition holds depends on the upper tail of match quality. Extreme-value theory yields tail-class loss laws with a sharp dichotomy: for light or bounded tails the loss vanishes as the candidate pool grows, whereas under heavy tails it converges to a size-independent constant, so expanding the candidate pool, even by orders of magnitude, does not asymptotically eliminate the cost. Evidence from two production experiments on a platform with millions of active creators is consistent with this picture: a pure A/A traffic sweep reveals a measurable, depth-graded engagement cost; a one-sided catalog ablation independently shows that per-viewer thinning contributes to the loss; and a tail index calibrated on the small exploration pool predicts an effect consistent with the one observed in the far larger full-catalog ablation. Isolation thus carries a price that experimenters should budget for, like any other cost. We give practitioners a preflight procedure that estimates it before launch, sizes traffic accordingly, and recommends a fallback design when the predicted cost exceeds a chosen tolerance.
Discretization and Statistical Consistency of Functional Flow Matching
oai:arXiv.org:2608.04531v1
arXiv:2608.04531v1 Announce Type: cross
Abstract: Functional flow matching is posed on distributions of functions but implemented from finitely many coefficients or point values. Under scattered or adaptive refinement, the resulting conditioning sigma-algebras need not be nested, so martingale convergence does not justify the sensor limit. We prove strong $L^2$ convergence of finite conditional velocity targets for every strongly consistent sequence of finite-rank reconstructions, with quantitative bounds for orthogonal projections and a point-sensor extension through a regularity space. For learned flows, coupling directly to a population superposition path yields an end-to-end Wasserstein bound without assuming uniqueness of the population finite-dimensional ODE. We verify sensor-independent constants for a normalized quadrature neural operator, including globally Lipschitz activations through an explicit magnitude recurrence. A noncommuting trace-class Gaussian example gives boundary multiplier $0$ under projected restriction and $0.72$ under exact conditioning. A spatial regularity--cubature certificate closes the operator-realization term, a Bernstein argument gives a $\widetilde{O}(n^{-1})$ excess-risk term for fixed model dimension and envelopes, and an exactly realizable clipped Gaussian scaling specialization yields an explicit end-to-end rate.
An entropic explanation of insistence on sameness in autism
oai:arXiv.org:2608.04616v1
arXiv:2608.04616v1 Announce Type: cross
Abstract: An information theory-based framework is proposed in attempt to explain insistence on sameness in autism as an instance of a general behavior pattern in which an individual tries to reduce surprise and uncertainty. It offers a new definition of autism as an impairment in which cognitive functions are restricted to discrimination, memorization and prediction of tangible properties of the environment. An analogy between insistence on sameness and constrained minimization of the entropy metric is observed and examined for a set of assumptions that describe cognitive limitations of a person with autism. The metric is given by the formula $D_H(R, M) = H(R|M) + H(M|R)$, where $R$ represents sequences of random stimuli, $M$ is a memory that stores and retrieves them, and where $H(.|.)$ denotes their conditional entropies interpreted as surprise and uncertainty, respectively. It is first inferred that to minimize the metric an individual can learn about $R$ (and store that knowledge in $M$) or can restrict $R$ to the already known $M$. Then, it is concluded that insistence on sameness is a manifestation of the latter. Moreover, it is shown that the proposed framework: (1) Helps to quantify the concepts of surprise, uncertainty, sensory overload and deprivation, anxiety, comfort zone, disappointment, disorientation, pedantry, rigidness, observance or aberrant precision. (2) Leads to a list of guidelines for learning therapies and daily care routines, and allows them to be defined as optimization algorithms and implemented as programs for robotic live-in caregivers. (3) Can be validated with the help of a Turing test-like approach that requires no experiments involving individuals with autism. The framework-if positively validated-will provide formal foundations and design guidelines for therapies aimed at improving self-reliance of individuals with autism in basic activities of daily living.
Drivers of Success: A Bayesian State-Space Model to Disentangling Latent Driver and Constructor Abilities in Formula One
oai:arXiv.org:2608.04629v1
arXiv:2608.04629v1 Announce Type: cross
Abstract: Formula One outcomes reflect the joint contributions of drivers and constructors, but these contributions are unobserved and vary over time. We propose a Bayesian state-space model that disentangles dynamic driver and constructor abilities using two observed outcomes: fastest qualifying lap times and race rankings. Both outcomes depend jointly on latent driver and constructor states that evolve at the Grand Prix level, while the race equation additionally accounts for starting-grid position. The decomposition is supported by constraints that center the driver and constructor abilities at zero, together with variation in driver-constructor assignments over time. Bayesian inference is performed using the No-U-Turn sampler under weakly informative priors that treat driver and constructor abilities symmetrically. Applying the model to the Formula One hybrid era from 2014 to 2021, we find substantial heterogeneity in both driver and constructor abilities. Driver abilities are generally more stable over time, whereas constructor abilities exhibit greater variation and, for many driver--constructor combinations, contribute more strongly to observed performance.
Clustered Local Projections for Short and Ultra-Short Time Series -- A Hierarchical Bayesian Framework
oai:arXiv.org:2608.04631v1
arXiv:2608.04631v1 Announce Type: cross
Abstract: Estimating the dynamic effects of economic shocks in short and very short samples is impeded by a lack of degrees of freedom. We offer a solution based on a Bayesian hierarchical framework for estimating local projection (LP) impulse response functions across a panel of related time series. The framework explicitly accommodates unbalanced panels in which some series are substantially shorter than others, allowing the short series to borrow information from longer ones at horizons where the short series carry little or no own data. Since series might exhibit heterogeneous dynamics, we develop a sparse finite mixture pool that clusters units by similarity of their impulse response profiles. We show in simulations that our approach substantially improves LP estimation accuracy relative to the standard approach if the time series are short while producing similar LPs for longer time series. Using a US price dataset, augmented with survey responses, we find that supply-chain and oil shocks trigger heterogeneous reactions of different price measures, with headline price indices responding more sharply than their core counterparts and goods prices changing more than services prices.
Personalized Federated Sparse Adaptation of Time-Series Foundation Models
oai:arXiv.org:2608.04695v1
arXiv:2608.04695v1 Announce Type: cross
Abstract: Federated adaptation of time-series foundation models (TSFMs) is attractive for building energy forecasting because meter data are private, distributed, and highly non-IID. However, a single parameter-sharing strategy is unlikely to serve all pretrained TSFMs or building clients: fully shared adapters can suppress building-specific temporal behavior, while fully local adaptation discards cross-building transfer. We propose a personalized federated sparse adaptation framework with a heterogeneous temporal mixture-of-experts (MoE) adapter placed after the pretrained TSFM representation. A sequence-level router maps each 168-hour context window to a top-$k$ subset of experts specialized for periodicity, long-range interactions, local variation, trend-residual structure, and multi-resolution behavior. We compare global FL, local training, and personalized FL variants with globally shared or client-private expert banks. Across 50 buildings and three TSFM backbones, personalization consistently outperforms Global FL-MoE and Local MoE, while the best sparse-adaptation strategy varies by backbone and metric. Routing behavior further reveals client-level expert specialization, expert concentration, and near-uniform routing across backbones, showing that federated TSFM adaptation should be both client-aware and backbone-aware.
Strong Convergence for a General Class of Random Matrix Models
oai:arXiv.org:2608.04824v1
arXiv:2608.04824v1 Announce Type: cross
Abstract: Let \(X_{1,n},\ldots,X_{d,n}\) be \(n\times n\) random matrices built from independent i.i.d. entry arrays, with centered entries, normalized by \(n^{-1/2}\). We prove that, if every entry law has finite fourth moment, then this tuple converges almost surely strongly in \(*\)-distribution to a free circular family with the matching variances. Equivalently, normalized traces and operator norms converge for every fixed noncommutative \(*\)-polynomial, including polynomials with fixed matrix coefficients. No assumption is imposed on the pseudo-variances of the complex entries. The bounded-entry argument applies the spectrum and moment universality estimates of Brailovskaya and van Handel to all self-adjoint linear pencils. The matching Gaussian pencils are reduced to independent Wigner matrices and identified by Anderson's strong convergence theorem. A fixed-level centered truncation, followed by the Bai--Yin norm bound, transfers the result to finite fourth moments.
Variational Bounds for Perceptron Learning from Structured Data
oai:arXiv.org:2608.04882v1
arXiv:2608.04882v1 Announce Type: cross
Abstract: We introduce a variational approach to a finite-temperature continuous-spin perceptron trained on a Gaussian mixture. The model allows for a broad class of concave utilities and log-concave separable prior measures on the spins. By combining the interpolation method with log-concavity and concentration estimates, we derive lower and upper minimax variational bounds for the limiting quenched pressure. Remarkably, the two bounds differ only in the order of optimization of two variational parameters, while all remaining extrema are controlled by the concave--convex structure of the variational potential. Whenever the two optimizations commute, the two bounds match and identify the solution of the model. The same potential yields the fixed-point equations as stationarity conditions and provides a unified route to the computation of the ground-state energy, training loss, and generalization error.
A Sharper Hoeffding Bound for Weighted Sums of Exchangeable Random Variables
oai:arXiv.org:2608.04900v1
arXiv:2608.04900v1 Announce Type: cross
Abstract: We prove a Hoeffding-type moment generating function bound for weighted sums of bounded exchangeable random variables centered by their finite-population average. The bound improves the finite-population inflation factor in a recent weighted exchangeable Hoeffding inequality from logarithmic order to the rate-optimal inverse-population-size order, with an explicit constant. The proof reduces the problem to Hamming slices, identifies two-level extremizers for the relevant symmetric variational problem, and applies a hypergeometric martingale bound. We also give a lower bound showing that an inverse-population-size inflation is unavoidable.
Stochastic Emulation using Generalized Stratified Sampling for Performance-Based Risk Optimization of Structures
oai:arXiv.org:2608.05006v1
arXiv:2608.05006v1 Announce Type: cross
Abstract: Metamodels are instrumental in reducing the computational burden associated with nested reliability analyses and optimization loops in Performance-Based Risk Optimization (PBRO) of structures under stochastic loads. In this context, stochastic emulators are particularly useful because they approximate response distributions while accounting for the intrinsic stochasticity of the simulator. Among these methods, Stochastic Polynomial Chaos Expansion (SPCE) is especially attractive because it does not require replications of nonlinear analyses at fixed input conditions. However, SPCE may present limitations in accurately representing extreme responses in the tails of structural response distributions. To address this limitation, this study proposes a framework that combines Generalized Stratified Sampling (GSS) with SPCE. The GSS scheme partitions the input space into strata according to the intensity of the hazard, improving the representation of extreme responses, while independent SPCE emulators are trained within each stratum. The conditional exceedance probabilities estimated in each stratum are then recombined using the total probability theorem to evaluate the probabilistic constraints. The proposed GSS-SPCE framework is applied to the optimal design of buckling-restrained brace cross-sectional areas in a two-story steel building. The objective is to minimize the initial construction cost while satisfying prescribed probabilistic performance constraints. Results show that the proposed framework accurately estimates structural response distributions, including their tail regions, while substantially reducing the number of nonlinear model evaluations required for PBRO.
Algorithm-Driven SVARs: Navigating the Wilderness of Big Data
oai:arXiv.org:2608.05017v1
arXiv:2608.05017v1 Announce Type: cross
Abstract: Every SVAR result is conditional on two choices: the restrictions that identify the shock and the variables on which they operate. The literature disciplines the first; the second is chosen by hand. We develop a Bayesian methodology that constructs information sets, uses an out-of-sample criterion, and retains the largest system it admits. Under recursive identification, output rises with housing production rather than household credit alone. For monetary policy, an anchor-free joint Bayesian proxy SVAR with multiple instruments strengthens the credit spread channel. A core system augmented with the selected corporate spread identifies expected default risk as a potent transmission margin.
Exact simulation of diffusions and improved algorithms for log-concave sampling
oai:arXiv.org:2608.05022v1
arXiv:2608.05022v1 Announce Type: cross
Abstract: We study exact simulation of diffusions via rejection sampling on path space using unbiased estimators of the density ratio obtained from Girsanov's theorem. When applied to the underdamped Langevin diffusion, it yields an algorithm for sampling from a strongly log-concave and log-smooth distribution with condition number $\kappa$, in dimension $d$, to accuracy $\varepsilon$ in R\'enyi divergence, in $\widetilde O(\kappa^{2/3} d^{1/3}\,\mathrm{polylog}(1/\varepsilon))$ queries. Under a third derivative bound, the dimension dependence improves to $d^{1/5}$. This improves substantially over the prior state-of-the-art complexity of $\widetilde O(\kappa d^{1/2}\,\mathrm{polylog}(1/\varepsilon))$ for the Metropolis-adjusted Langevin algorithm, and over the $d^{1/4}$ dimension dependence of Metropolized Hamiltonian Monte Carlo under the same third derivative bound. We also present applications to the mirror Langevin diffusion, and for obtaining Fisher information bounds in the non-log-concave case.
Canonical Joint Energy-Based Model on CIFAR-10: failure modes and practical indistinguishability of Predictor-Corrector and SGLD samplers
oai:arXiv.org:2608.05025v1
arXiv:2608.05025v1 Announce Type: cross
Abstract: Joint Energy-Based Models (JEM) unify classification and generation within a single network and support out-of-distribution (OOD) detection. Canonical JEM training relies on stochastic gradient Langevin dynamics (SGLD); a theoretically motivated alternative, the Predictor-Corrector (PC) sampler, has not previously undergone a systematic replication test on the canonical model. We reproduce canonical JEM on WideResNet-28-10 without normalisation layers on two independent runs and test whether PC retains its theoretical advantage without an annealed noise schedule, across three protocols: PC replacing SGLD throughout the roughly 130 training epochs; cold-start generation (FID); and refinement-style multi-OOD detection (AUROC). The reconstruction reaches 92.88% test accuracy and buffer-FID 44.46 (canonical: 92.9% and 38.40). We document two failure modes: catastrophic late-training divergence via the canonical outlier-buffer mechanism (both SGLD runs and, with the same signature, both PC runs), and run-dependent SVHN OOD-discrimination dynamics. No method-level advantage of PC over SGLD is observed on any protocol: at inference the absolute AUROC difference stays below 0.007 across all ten checkpoint-OOD pairs and the FID difference below 0.5; on the training protocol a hierarchical seed-by-image bootstrap gives a 95% confidence interval on the macro-averaged AUROC difference that contains zero, while a seed-level equivalence test with two runs per method cannot establish formal equivalence. The data are consistent both with equivalence and with a small directional effect. This practical indistinguishability is theoretically expected: under fixed noise the PC predictor step degenerates by construction, so its guarantees do not transfer to canonical JEM.
Representational separation between unitary and channel quantum generative models via shared classical randomness at shallow depth
oai:arXiv.org:2608.05110v1
arXiv:2608.05110v1 Announce Type: cross
Abstract: Near-term quantum hardware limits circuit depth and often imposes geometrically local connectivity for quantum generative models, restricting the output distributions accessible to shallow unitary Born models. Introducing stochasticity into a unitary quantum Born model can improve the empirical generative performance of the resulting channel model and, for a restricted small-scale architecture, has been proven to represent a strictly larger family of distributions than its unitary counterpart. However, whether such randomness provides a provable separation at fixed shallow depth for arbitrarily large systems has remained open. Here, we show that shared classical randomness, a comparatively weak resource from entanglement theory, is sufficient to establish such a strict scalable representational separation over the corresponding shallow unitary Born model. More specifically, we augment bounded-connectivity shallow unitary circuits, followed by computational-basis measurements, with spatially separated local Pauli operations, whose joint application is controlled by a single classically sampled random bit. The resulting shallow-depth channel model generates long-range correlations in the classical output distribution that no purely unitary shallow-depth model with bounded connectivity can reproduce. For one-dimensional nearest-neighbour architectures, reproducing such distributions with a purely unitary model can require depth $\Omega(N)$ in the worst case. We further show that measurement-based quantum computation (MBQC) provides a natural implementation of the required shared classical randomness through suitable adaptation of the random measurement outcomes. Numerical experiments on MBQC-based generative models support the analytical results.
SSTQ:Privacy-Preserving Vector Quantization via Subsampled Stochastic TurboQuant
oai:arXiv.org:2608.05127v1
arXiv:2608.05127v1 Announce Type: cross
Abstract: Achieving local differential privacy in distributed optimization while maintaining low communication cost remains challenging. Existing vector quantization methods, such as vqSGD, use high-dimensional geometric constructions but incur unfavorable dimension-dependent variance. In this work, we propose Subsampled Stochastic TurboQuant (SSTQ), a framework that combines overcomplete equal-norm tight frames, coordinate subsampling, and privacy-aware one-dimensional quantization. SSTQ includes two variants: a Flat Randomized Response version and a Metric-Aware Laplace version, the latter being better suited to higher codebook bit-width regimes. We show that SSTQ achieves optimal mean squared error scaling while using only $\lceil \log_2 N \rceil + b$ bits per client, where $N = \Theta(d)$ is the frame size. We also derive a surrogate privacy-aware codebook objective that reduces the codebook-dependent MSE scaling from $O(4^b)$ to $O(2^b)$. Finally, we empirically evaluate SSTQ against established baselines on federated learning tasks using CIFAR-10 and Fashion-MNIST, demonstrating favorable utility and communication efficiency.
Partial Identification of Causal Effects Using Proxy Variables
oai:arXiv.org:2304.04374v4
arXiv:2304.04374v4 Announce Type: replace
Abstract: Proximal causal inference is a framework for evaluating the causal effects in the presence of unmeasured confounding. For point identification, it leverages a pair of proxy variables to identify a bridge function that matches the dependence of potential outcomes or treatment variables on the hidden factors to corresponding functions of observed proxies. Unique identification requires that proxies are sufficiently relevant for hidden factors, a requirement that has previously been formalized as a completeness condition. However, completeness is not empirically testable, and although a bridge function may be well-defined in a given setting, lack of completeness, sometimes manifested by availability of a single type of proxy, may severely limit prospects for identification of a bridge function and thus a causal effect; therefore, potentially restricting the application of the framework. In this paper, we propose partial identification methods that do not require completeness and obviate the need for identification of a bridge function. We establish that proxies can be leveraged to obtain bounds on the causal effect even if available information does not suffice to identify either a bridge function or a corresponding causal effect of interest. Our bounds are non-smooth functionals of the underlying distribution. For inference, we employ LogSumExp approximations that yield smooth lower and upper bounds, and we derive the efficient influence functions of the resulting bound functionals which enable analytic variance estimation, while bootstrap confidence intervals remain available for regular plug-in implementations. We further establish analogous results in related settings where identification hinges upon hidden mediators for which proxies are available, however such proxies are not sufficiently rich for point identification of a bridge function or a corresponding causal effect of interest.
E$^2$M: Double Bounded $\alpha$-Divergence Optimization for Tensor-based Discrete Density Estimation
oai:arXiv.org:2405.18220v4
arXiv:2405.18220v4 Announce Type: replace
Abstract: Tensor-based discrete density estimation requires flexible modeling and proper divergence criteria to enable effective learning; however, traditional approaches using $\alpha$-divergence face analytical challenges due to the $\alpha$-power terms in the objective function, which hinder the derivation of closed-form update rules. We present a generalization of the expectation-maximization (EM) algorithm, called the E$^2$M algorithm. It circumvents this issue by first relaxing the optimization into the minimization of a surrogate objective based on the Kullback-Leibler (KL) divergence, which is tractable via the standard EM algorithm, and subsequently applying a tensor many-body approximation in the M-step to enable simultaneous closed-form updates of all parameters. Our approach offers flexible modeling for not only a variety of low-rank structures, including the CP, Tucker, and Tensor Train formats, but also their mixtures, thus allowing us to leverage the strengths of different low-rank structures. We evaluate the effectiveness of our approach on synthetic and real datasets, highlighting its comparable convergence to gradient-based procedures, robustness to outliers, and favorable density estimation performance compared to prominent existing tensor-based methods.
How accurate are Bayes factor-based null hypothesis tests? A simulation study
oai:arXiv.org:2406.08022v3
arXiv:2406.08022v3 Announce Type: replace
Abstract: Bayes factor null hypothesis tests provide a viable alternative to frequentist measures of evidence quantification. Bayes factors for realistic data sets in areas like psychology cannot be calculated exactly and require numerical approximations to complex integrals. Crucially, the accuracy of these approximations, i.e., whether an approximate Bayes factor corresponds to the exact Bayes factor, is unknown, and may depend on data, prior, and likelihood. We have recently developed a novel statistical procedure, namely marginal simulation-based calibration (SBC) for Bayes factors, to test whether the computed Bayes factors for a given analysis are accurate. Here, we use marginal SBC for Bayes factors and calibration plots to test for some common cognitive designs, whether Bayes factors are calculated accurately. We use the bridgesampling/brms packages in R. We run analyses for three commonly used designs in psychology and psycholinguistics: (a) a design with random effects for subjects only, (b) a Latin square design with crossed random effects for subjects and items, but a single fixed-factor, and (c) a Latin square 2x2 design with crossed random effects for subjects and items. We find that Bayes factor estimates turn out accurate in cases when the bridgesampling algorithm does not issue a warning message, but can be biased and variable when a warning message is shown. These results support the use of brms/bridgesampling for null hypothesis Bayes factor tests in commonly used factorial designs. They also suggest that when a warning message is issued, Bayes factor results should not be trusted. The results show that it is practical to check whether Bayes factors are computed correctly.
Non-asymptotic Estimates for Markov Transition Matrices via Spectral Gap Methods
oai:arXiv.org:2408.05963v4
arXiv:2408.05963v4 Announce Type: replace
Abstract: We establish non-asymptotic error bounds for the classical Maximal Likelihood Estimation of the transition matrix of a given Markov chain. Meanwhile, in the reversible case, we propose a new reversibility-preserving online Symmetric Counting Estimation of the transition matrix with non-asymptotic deviation bounds. Our analysis is based on a convergence study of certain Markov chains on the length-2 path spaces induced by the original Markov chain.
Aspects of a Generalized Theory of Sparsity based Inference in Linear Inverse Problems
oai:arXiv.org:2503.00178v2
arXiv:2503.00178v2 Announce Type: replace
Abstract: Linear inverse problems are ubiquitous in various science and engineering disciplines. Of particular importance in the past few decades, is the incorporation of sparsity based priors, in particular $\ell_1$ priors, into linear inverse problems, which led to the flowering of fields of compressive sensing (CS) and sparsity based signal processing. More recently, methods based on a Compound Gaussian (CG) prior have been investigated and demonstrate improved results over CS in practice. This paper is the first attempt to identify and elucidate the fundamental structures underlying the success of CG methods by studying CG in the context of a broader framework of generalized-sparsity-based-inference. After defining our notion of generalized sparsity we introduce a weak null space property and proceed to generalize two well-known methods in CS, basis pursuit and iteratively reweighted least squares (IRLS). We show how a subset of CG-induced regularizers fits into this framework.
MoCA: Multi-modal Cross-masked Autoencoder for Digital Health Measurements
oai:arXiv.org:2506.02260v4
arXiv:2506.02260v4 Announce Type: replace
Abstract: Wearable devices enable continuous multi-modal physiological and behavioral monitoring, yet analysis of these data streams faces fundamental challenges including the lack of gold-standard labels and incomplete sensor data. While self-supervised learning approaches have shown promise for addressing these issues, existing multi-modal extensions present opportunities to better leverage the rich temporal and cross-modal correlations inherent in simultaneously recorded wearable sensor data. We propose the Multi-modal Cross-masked Autoencoder (MoCA), a self-supervised learning framework that combines transformer architecture with masked autoencoder (MAE) methodology, using a principled cross-modality masking scheme that explicitly leverages correlation structures between sensor modalities. MoCA demonstrates strong performance boosts across reconstruction and downstream classification tasks on diverse benchmark datasets. We further establish theoretical guarantees by establishing a fundamental connection between multi-modal MAE loss and kernelized canonical correlation analysis through a Reproducing Kernel Hilbert Space framework, providing principled guidance for correlation-aware masking strategy design. Our approach offers a novel solution for leveraging unlabeled multi-modal wearable data while handling missing modalities, with broad applications across digital health domains.
Z-Curve Plot: A Visual Diagnostic for Publication Bias in Meta-Analysis
oai:arXiv.org:2509.07171v2
arXiv:2509.07171v2 Announce Type: replace
Abstract: Publication bias undermines meta-analytic inference, yet visual diagnostics for detecting and understanding model misfit due to publication bias are lacking. We propose the z-plot, a visual publication bias-focused absolute model fit diagnostic. The z-plot overlays the model-implied distribution of z-statistics on the observed distribution of z-statistics. Models that approximate the data well show minimal discrepancy between the observed and predicted distributions of z-statistics, whereas models that approximate the data poorly show systematic discrepancies. Discontinuities in the observed distribution of z-statistics at significance thresholds or at zero provide visual evidence of publication bias; models that account for this bias track these discontinuities. In addition, the z-plot facilitates visual model fit comparison of competing meta-analytic models within a single figure. We demonstrate the visualization and its interpretation with a Bayesian random-effects meta-analysis, a Bayesian PET model, a Bayesian three-parameter selection model, and RoBMA on simulated datasets and a real meta-analysis. The method is implemented in the RoBMA R package.
Meta-Analysis with JASP, Part I: Classical Approaches
oai:arXiv.org:2509.09845v2
arXiv:2509.09845v2 Announce Type: replace
Abstract: Meta-analyses play a crucial part in empirical science, enabling researchers to synthesize evidence across studies and draw more precise and generalizable conclusions. Despite their importance, access to advanced meta-analytic methodology is often limited to scientists and students with considerable expertise in computer programming. To lower the barrier for adoption, we have developed the Meta-Analysis module in JASP (https://jasp-stats.org/), a free and open-source software for statistical analyses. The module offers standard and advanced meta-analytic techniques through an easy-to-use graphical user interface (GUI), allowing researchers with diverse technical backgrounds to conduct state-of-the-art analyses. This manuscript presents an overview of the meta-analytic tools implemented in the module and showcases how JASP supports a meta-analytic practice that is rigorous, relevant, and reproducible. Tutorial videos accompany the examples presented in this manuscript.
Convergence and Stability Analysis of Self-Consuming Generative Models with Heterogeneous Human Curation
oai:arXiv.org:2511.09002v3
arXiv:2511.09002v3 Announce Type: replace
Abstract: Self-consuming generative models have received significant attention over the last few years. In this paper, we study a self-consuming generative model with heterogeneous preferences that is a generalization of the model in Ferbach et al. (2024). The model is retrained round by round using real data and its previous-round synthetic outputs. The asymptotic behavior of the retraining dynamics is investigated across four regimes using different techniques including the nonlinear Perron--Frobenius theory. Our analyses improve upon that of Ferbach et al. (2024) and provide convergence results in settings where the well-known Banach contraction mapping arguments do not apply. Stability and non-stability results regarding the retraining dynamics are also given.
Differentially Private Tests of Fisher's Sharp Null Hypothesis for Binary Outcomes
oai:arXiv.org:2511.20884v3
arXiv:2511.20884v3 Announce Type: replace
Abstract: Across many disciplines, causal inference often relies on randomized experiments with binary outcomes. In such experiments, often analysts are interested in testing Fisher's sharp null hypothesis, i.e., the treatment has zero effect on the outcome for every study subject. Sometimes the outcomes are sensitive and must be kept confidential, for example, when they comprise physical or mental health measurements. Releasing test statistics or p-values computed with the confidential outcomes can leak information about the individuals in the study. Those responsible for sharing the analysis results may wish to bound this information leakage, which they can do by ensuring the released outputs satisfy differential privacy. In this article, we develop several differentially private tests of Fisher's sharp null for binary outcomes. Specifically, we consider direct perturbation approaches that inject calibrated noise into test statistics or p-values, as well as a Bayesian denoising framework that explicitly models the privacy mechanism. We further develop decision-making procedures under privacy constraints, including a Bayes risk-optimal rule and a frequentist-calibrated significance test. Through theoretical results, simulation studies, and an application to the ADAPTABLE clinical trial, we demonstrate that our methods can achieve valid and interpretable causal inference while ensuring the differential privacy guarantee.
Predicting Dry Spells of the West African Monsoon Season Using Machine Learning Methods
oai:arXiv.org:2512.01965v3
arXiv:2512.01965v3 Announce Type: replace
Abstract: Characteristics of the West African Monsoon (WAM) season, such as its onset and dry spell occurrences, are notoriously difficult to predict. However, these characteristics are key indicators farmers use to decide when to plant crops, having a major influence on their overall yield. While many studies have shown correlations between global sea surface temperatures and characteristics of the WAM season, there are few that effectively implement this information into machine learning (ML) prediction models. This study is focused on predicting dry spells, that is, if there will be a period of consecutive days without rain after the onset of the WAM. We first investigated the best ways to define onset and dry spells and gathered sea surface temperature training data from both real-world observations and a climate simulation model. Then we constructed an adaptive-threshold logistic regression model for dry spell prediction, to which we applied a custom feature selection method and spatial regularization. Using Leave-One-Out cross validation testing, we found significant results in multiple binary classification metrics. These models overcome some limitations that current approaches have, such as being computationally intensive and needing bias correction. We also aim for this study to serve as a framework for ML use in the context of targeted prediction of certain weather phenomena using climatologically relevant variables.
Energy-Tweedie: Score meets Score, Energy meets Energy
oai:arXiv.org:2512.23818v2
arXiv:2512.23818v2 Announce Type: replace
Abstract: Denoising and score estimation are classically linked through Tweedie's formula, which relates the posterior mean under Gaussian noise to the Stein score of the noisy marginal. In this work, we extend this perspective beyond Gaussian noise to a broad class of Gibbs (energy-based) noise distributions, with the generalized Gaussian family as the running example. We derive the Energy-Tweedie identity: when the denoising posterior is viewed through the lens of scoring rules, the path derivative of a kernel scoring rule defined by the noise potential recovers the Stein score of the noisy marginal. The rule's propriety is determined by the noise potential alone. Thus, the familiar correspondence between Gaussian noise, posterior means, squared loss, and Tweedie's formula is lifted to a distributional correspondence between Gibbs noise distributions, full posterior laws, kernel scoring rules, and the Energy-Tweedie identity, yielding one Tweedie-style relation for each noise potential. Among its consequences, this identity gives a posterior-samples-to-score route to score estimation, yields a principled criterion for estimating unknown noise parameters, and enables diffusion-style sampling along user-chosen paths through the noise-parameter space, supplying the score-based perspective on recent generative methods trained with scoring rules.
Local Asymptotic Normality for Mixed Fractional Brownian Motion with $0oai:arXiv.org:2512.24042v2
arXiv:2512.24042v2 Announce Type: replace
Abstract: This paper establishes the Local Asymptotic Normality (LAN) property for the mixed fractional Brownian motion under high-frequency observations with Hurst index $H \in (0, 3/4)$. The simultaneous estimation of the volatility and the Hurst index encounters a degeneracy problem in the Fisher information matrix.
Wedge Sampling: Efficient Tensor Completion with Nearly-Linear Sample Complexity
oai:arXiv.org:2602.05869v3
arXiv:2602.05869v3 Announce Type: replace
Abstract: We introduce Wedge Sampling, a new non-adaptive sampling scheme for low-rank tensor completion. We study recovery of an order-$k$ low-rank tensor of dimension $n\times\cdots\times n$ from structured observations of its entries. Unlike the standard uniform entry model (i.e., i.i.d. samples from $[n]^k$), wedge sampling allocates observations to structured length-two patterns (wedges) in an associated bipartite sampling graph. By directly promoting these length-two connections, the sampling design strengthens the spectral signal that underlies efficient initialization, in regimes where uniform sampling is too sparse to generate enough informative correlations.
Our main result shows that this change in sampling paradigm enables polynomial-time algorithms to achieve both weak and exact recovery with nearly linear sample complexity in $n$. The approach is also plug-and-play: wedge-sampling-based spectral initialization can be combined with existing refinement procedures (e.g., spectral or gradient-based methods) using only an additional $\tilde O(n)$ uniformly sampled entries, substantially improving over the $\tilde O(n^{k/2})$ sample complexity typically required under uniform entry sampling for efficient methods. We also formulate a noisy wedge-sampling extension for additive Gaussian observations and analyze both the spectral and gradient-descent procedures under suitable signal-to-noise conditions. Thus, the computational barrier in tensor completion is sensitive to the observation model: while it persists under uniform entry sampling, it can be bypassed by non-adaptive structured designs that provide a stronger initialization.
Structured Semiparametric Estimation of Average Treatment Effects with Treatment-Specific Non-Gaussian Error Distributions
oai:arXiv.org:2604.07770v3
arXiv:2604.07770v3 Announce Type: replace
Abstract: This paper studies average treatment effect (ATE) estimation for continuous outcomes when the treatment-covariate mean is structured but the error distributions are unknown and may differ across treatment arms in scale, skewness, or tail behavior. We introduce a semiparametric model with a finite-dimensional mean, or a prespecified basis expansion, and separate unrestricted additive error laws for treated and control outcomes. The central theoretical contribution is the ATE-efficient influence function under two treatment-specific error nuisance spaces: it combines arm-specific regression-efficient scores with variation in the marginal covariate law and is distinct from both the unrestricted AIPW gradient and a pooled location-shift score. We establish the exact nested ordering of the efficiency bounds under pooled common-error, treatment-specific-error, and unrestricted causal models, including equality conditions. A local-misspecification result further shows that the variance reduction obtained by valid pooling equals the maximal squared first-order bias incurred along unit treatment-specific directions excluded by the pooled model. We develop efficient cross-fitted plug-in and ATE-targeted implementations requiring only two one-dimensional density estimates; the targeted refinement adds a scalar update. Simulations show substantial precision gains from valid pooling under common non-Gaussian errors and protection against invalid pooling when treatment changes error shape, including under weak positivity. Applications to ACTG175 and National Supported Work data illustrate the resulting precision--robustness tradeoff for irregular continuous outcomes.
Distributionally Robust Transfer Learning with Structurally Missing Covariates, with Application to Cross-National Cardiac Arrest Prediction
oai:arXiv.org:2605.24212v2
arXiv:2605.24212v2 Announce Type: replace
Abstract: Deploying clinical prediction models across healthcare systems often fails when key training covariates are unavailable at deployment and labeled outcomes are limited in the target domain. For example, high-performing models for out-of-hospital cardiac arrest (OHCA) rely on detailed prehospital measurements routinely collected in high-resource settings but unavailable in many international registries. Existing methods either discard missing covariates, sacrificing predictive information, or rely on untestable assumptions about their target distribution. We propose DRUM (\underline{D}istributionally \underline{R}obust \underline{U}nsupervised transfer learning with structurally \underline{M}issing covariates), a framework that transfers prediction models to target populations where certain covariates are structurally absent and outcome labels are unavailable. DRUM partitions covariates into shared components ($X$), observed across all settings, and missing components ($A$), observed only in the source. Rather than imputing missing covariates, DRUM optimizes worst-case predictive performance over the unknown target distribution of $A \mid X$ using a neural network generator, with a robustness parameter controlling allowable deviation from the source conditional. We further develop a bias correction procedure that reduces sensitivity to nuisance estimation error. Simulations show substantial improvements in both mean and worst-case prediction error under distribution shift. Applied to cross-national OHCA prediction, transferring models from a US registry to multiple Asian registries where prehospital variables are unrecorded, DRUM yields better-calibrated predictions and improved clinical classification performance across sites.
Sequential Kernel-based Conditional Independence Testing via Adaptive Betting
oai:arXiv.org:2606.18993v2
arXiv:2606.18993v2 Announce Type: replace
Abstract: Testing conditional independence is fundamental yet intrinsically difficult: without additional assumptions, Type I error control is impossible in general. The "Model-X'' paradigm addresses this difficulty by assuming exact knowledge of a relevant conditional distribution. While small deviations from this assumption can sometimes be tolerated in classical one-shot testing, existing sequential conditional independence tests typically require the Model-X conditional to be known exactly, making them fragile when it must instead be estimated. We propose a new approach that is substantially more robust to such estimation error. Our method applies testing-by-betting to an adaptively optimized Kernel Conditional Independence statistic, together with a normalization scheme and a truncate-and-shift calibration strategy. These modifications greatly reduce Type I error inflation while preserving high power across high-dimensional synthetic benchmarks and real-world fairness tasks, outperforming existing sequential Model-X approaches. Code is available at https://github.com/he-zh/SKCI.
The logistic-normal integral and the moments of the logistic-normal distribution
oai:arXiv.org:2607.07889v2
arXiv:2607.07889v2 Announce Type: replace
Abstract: The logistic-normal integral appears in problems of statistical estimation for logistic models with Gaussian random effects, and generalized linear mixed models. We study the numerical evaluation of this integral and of its derivatives, and give closed form evaluations at certain points and series expansions. There is a continuum of possible series expansions, and we single out one series expansion which is optimal for numerical evaluation. We propose an algorithm for a precise numerical evaluation, based on the optimal series, with good approximation error control in the tails. As an application we give explicit results for the first four moments of a logistic-normal random variable.
Weak Information Geometry: Riemannian Structures from Distributional Inference Functions and Stein Discrepancies
oai:arXiv.org:2607.11246v2
arXiv:2607.11246v2 Announce Type: replace
Abstract: The class of parametric statistical models that can be treated as Riemannian manifolds is considerably larger than the classical Fisher-Rao setting allows, once one works in the space of tempered distributions. A law is represented by a tempered distribution T in S'(R^k), while an instrument - a positive Schwartz kernel, a weak regular inference function, or a weak Stein representation - extracts information from the law without being part of it. Any instrument with full-rank sensitivity and positive-definite variability induces the Godambe information G = S^T V^{-1} S, a Riemannian metric on the parameter space; the Fisher-Rao manifold is recovered exactly when the score is an admissible instrument, and every Godambe metric is dominated by the Fisher metric in the Loewner order whenever the latter exists. Four examples lie outside the Fisher-Rao class for four different reasons: a location model built on the Cantor distribution (an undominated family - no likelihood, no score, and no Fisher information exist at all), the uniform scale model (parameter-dependent support), the shifted exponential model (transform-based inference), and a stratified finite mixture (a provably biased score in a dominated model); a lattice stochastic heat equation driven by alpha-stable noise provides a fifth, dynamical example, whose closed-form weak Godambe information stabilises at a rate governed by the spectral gap of the discrete Laplacian. Quadratic Stein discrepancies induce the same local geometry, and reproducing-kernel constructions generate a hierarchy of geometries. Because there is no canonical instrument, the model carries a family of Godambe metrics; we discuss the inferential, diagnostic, geometric, and computational roles of its members, and show that weak inferential separation (nonformation) appears geometrically as block-diagonality of the Godambe metric.
Mixing-Free and Signal-Optimal Learning of Gaussian Graphical Models from Glauber Dynamics
oai:arXiv.org:2607.18559v2
arXiv:2607.18559v2 Announce Type: replace
Abstract: Gaussian graphical model selection is usually studied under independent sampling, but in many applications the data arise as a single trajectory of a dependent stochastic process. We study exact recovery of the graph from one trajectory of random-scan Gaussian Glauber dynamics. Existing techniques for this problem either inherit the mixing time of the chain, which can be super-polynomial in the dimension $p$ without strong assumptions, or are suboptimal in the minimum normalized edge strength $\kappa$. We propose two algorithms that are mixing-free and attain the $\kappa^{-2}$ dependence of the information-theoretic lower bounds. Both instantiate a shared dueling-neighborhood search meta-algorithm with a local statistic built directly from the update sequence. For every fixed precision matrix and deterministic initialization, the first algorithm fits a least-squares regression at the updates of each node and has pointwise recovery horizon $\widetilde O(pd^{2}/\kappa^{2})$, where $d$ is the maximum degree. Its horizon depends logarithmically on a local conditioning quantity and on the initialization potential. The second algorithm is based on counting occurences of a specific update pattern and requires $\widetilde O(pd^{4}/\kappa^{2})$ updates, with no dependence on any condition number. The central technical challenge is that both statistics are built from dependent, non-stationary observations. Our analysis tackles this by demonstrating how to extract fresh Gaussian innovations from the update sequence, which yields mixing-free control of appropriate quantities. Neither the algorithms nor their analyses invoke stationarity, a spectral gap, or mixing conditions.
Topological Clustering via Sliced Wasserstien Kernels
oai:arXiv.org:2608.03891v2
arXiv:2608.03891v2 Announce Type: replace
Abstract: Topological data analysis (TDA) uses topological techniques to extract meaningful shape-based information from complex datasets. Clustering is a central problem in data analysis, and there has been considerable recent interest in understanding how TDA can inform it. Existing approaches either cluster persistence diagrams directly under Wasserstein-type distances, which is computationally expensive, or use vector representations of diagrams. We propose a kernel $k$-means algorithm built on a convex combination of sliced Wasserstein (SW) kernels, one for each homology under consideration. Unlike other vector representations of persistence diagrams, the SW kernel is both stable and discriminative with respect to the $1$-Wasserstein distance. The method outperforms the baselines on two benchmark datasets and remains competitive on a third synthetic dataset, while being computationally efficient. It also outperforms both a single SW kernel on the union of all homology groups and an SW kernel computed directly on the point clouds. The convex combination assigns an interpretable weight to each $q$-homology kernel. We further validate that the weights identify the discriminating homology.
Robust Low-Tubal-Rank Tensor Completion under Cross-Concentrated Sampling
oai:arXiv.org:2608.03928v2
arXiv:2608.03928v2 Announce Type: replace
Abstract: Tensor cross-concentrated sampling (t-CCS) bridges entrywise sampling and t-CUR slice-wise sampling by observing entries only within selected horizontal and lateral slices. Existing t-CCS completion methods, however, assume that the observations are free of gross corruption. In this work, we study robust recovery of a third-order low-tubal-rank tensor from partial t-CCS observations contaminated by sparse, arbitrarily large outliers. We propose Robust Iterative t-CUR (R-ItCUR), a tensor-native algorithm that partitions the sampled tensor cross into two exterior blocks and an intersection block, applies adaptive blockwise Welsch correction for outlier suppression, and updates the low-rank component through projected blockwise gradient descent. By operating directly on the sampled cross, R-ItCUR avoids reconstructing the full tensor throughout the iterations, resulting in substantial memory and computational savings. Experiments on synthetic tensors, cardiac MRI data, and three-dimensional seismic data demonstrate accurate recovery and strong robustness to sparse gross corruptions. The results further highlight the importance of explicitly exploiting the cross-concentrated sampling structure in robust tensor completion.
GFlowNet Training by Policy Gradients
oai:arXiv.org:2408.05885v3
arXiv:2408.05885v3 Announce Type: replace-cross
Abstract: Generative Flow Networks (GFlowNets) have been shown effective to generate combinatorial objects with desired properties. We here propose a new GFlowNet training framework, with policy-dependent rewards, that bridges keeping flow balance of GFlowNets to optimizing the expected accumulated reward in traditional Reinforcement-Learning (RL). This enables the derivation of new policy-based GFlowNet training methods, in contrast to existing ones resembling value-based RL. It is known that the design of backward policies in GFlowNet training affects efficiency. We further develop a coupled training strategy that jointly solves GFlowNet forward policy training and backward policy design. Performance analysis is provided with a theoretical guarantee of our policy-based GFlowNet training. Experiments on both simulated and real-world datasets verify that our policy-based strategies provide advanced RL perspectives for robust gradient estimation to improve GFlowNet performance.
Regularization can make diffusion models more efficient
oai:arXiv.org:2502.09151v3
arXiv:2502.09151v3 Announce Type: replace-cross
Abstract: Diffusion models are one of the key architectures of generative AI. Their main drawback, however, is the computational costs. This study indicates that the concept of sparsity, well known especially in statistics, can provide a pathway to more efficient diffusion pipelines. Our mathematical guarantees prove that sparsity can reduce the input dimension's influence on the computational complexity to that of a much smaller intrinsic dimension of the data. Our empirical findings confirm that inducing sparsity can indeed lead to better samples at a lower cost.
Towards Understanding Gradient Flow Dynamics of Homogeneous Neural Networks Beyond the Origin
oai:arXiv.org:2502.15952v3
arXiv:2502.15952v3 Announce Type: replace-cross
Abstract: Recent works exploring the training dynamics of homogeneous neural network weights under gradient flow with small initialization have established that in the early stages of training, the weights remain small and near the origin, but converge in direction. Building on this, the current paper studies the gradient flow dynamics of homogeneous neural networks with locally Lipschitz gradients, after they escape the origin. Insights gained from this analysis are used to characterize the first saddle point encountered by gradient flow after escaping the origin. Also, it is shown that for homogeneous feed-forward neural networks, under certain conditions, the sparsity structure emerging among the weights before the escape is preserved after escaping the origin and until reaching the next saddle point.
GenAI-Powered Inference
oai:arXiv.org:2507.03897v3
arXiv:2507.03897v3 Announce Type: replace-cross
Abstract: We introduce GenAI-Powered Inference (GPI), a statistical framework for both causal and predictive inference using unstructured data, including text and images. GPI leverages open-source Generative Artificial Intelligence (GenAI) models---such as large language models and diffusion models---not only to generate unstructured data at scale but also to extract low-dimensional representations that are guaranteed to capture their underlying structure. Applying machine learning to these representations, GPI enables estimation of causal effects while quantifying associated estimation uncertainty. Unlike existing approaches to representation learning, GPI does not require fine-tuning of generative models, making it computationally efficient and broadly accessible. We illustrate the versatility of the GPI framework through three applications: (1) estimating the effects of Chinese social media censorship while adjusting for textual confounders, (2) isolating the impact of specific image features from that of other correlated features in the same image, and (3) assessing the persuasiveness of political rhetoric. An open-source software package is available for implementing GPI.
Learning Neural Networks by Neuron Pursuit
oai:arXiv.org:2509.12154v2
arXiv:2509.12154v2 Announce Type: replace-cross
Abstract: The first part of this paper studies the evolution of gradient flow for homogeneous neural networks near a class of saddle points exhibiting a sparsity structure. The choice of these saddle points is motivated from previous works on homogeneous networks, which identified the first saddle point encountered by gradient flow after escaping the origin. It is shown here that, when initialized sufficiently close to such saddle points, gradient flow remains near the saddle point for a sufficiently long time, during which the set of weights with small norm remain small but converge in direction. Furthermore, important empirical observations are made on the behavior of gradient descent after escaping these saddle points. The second part of the paper, motivated by these results, introduces a greedy algorithm to train deep neural networks called Neuron Pursuit (NP). It is an iterative procedure which alternates between expanding the network by adding neuron(s) with carefully chosen weights, and minimizing the training loss using this augmented network. The efficacy of the proposed algorithm is validated using numerical experiments.
Multicalibration Yields Better Matchings
oai:arXiv.org:2511.11413v2
arXiv:2511.11413v2 Announce Type: replace-cross
Abstract: Consider the problem of finding the best matching in a weighted graph where we only have access to predictions of the actual stochastic weights, based on an underlying context. If the predictor is the Bayes optimal one, then computing the best matching based on the predicted weights is optimal. However, in practice, this perfect information scenario is not realistic. Given an imperfect predictor, a suboptimal decision rule may compensate for the induced error and thus outperform the standard optimal rule.
In this paper, we propose multicalibration as a way to address this problem. This fairness notion requires a predictor to be unbiased on each element of a family of protected sets of contexts. Given a class of matching algorithms $\mathcal C$ and any predictor $\gamma$ of the edge-weights, we show how to construct a specific multicalibrated predictor $\hat \gamma$, with the following property. Picking the best matching based on the output of $\hat \gamma$ is competitive with the best decision rule in $\mathcal C$ applied onto the original predictor $\gamma$. We complement this result by providing sample complexity bounds, and by performing numerical experiments.
Bounds on inequality with incomplete data
oai:arXiv.org:2512.07709v3
arXiv:2512.07709v3 Announce Type: replace-cross
Abstract: We study inequality measures when outcomes are observed only in intervals, as in historical tabulations, privacy-protected grouped data, and modern surveys. We develop a nonparametric framework for sharp identification and inference with grouped and interval-valued data, covering brackets and overlapping intervals. For a class of inequality indices, sharp bounds are attained by discrete distributions with finite support, reducing the problem to optimization; linear-fractional indices, including the Gini and quantile ratios, yield linear or quadratic programs. Plug-in bound endpoints have a $\sqrt{n}$ asymptotic distribution, using an $m$-out-of-$n$ bootstrap. Applications to wealth and historical income data compare identified sets with imputation-based estimates.
Non-Stationary Inventory Control with Lead Times
oai:arXiv.org:2602.05799v2
arXiv:2602.05799v2 Announce Type: replace-cross
Abstract: We study non-stationary single-item, periodic-review inventory control problems in which the demand distribution is unknown and may change over time. We analyze how demand non-stationarity affects learning performance across inventory models, including systems with demand backlogging or lost-sales, both with and without lead times. For each setting, we propose an adaptive online algorithm that optimizes over the class of base-stock policies and establish performance guarantees in terms of dynamic regret relative to the optimal base-stock policy at each time step. The algorithms leverage the convexity and one-sided feedback structure of inventory costs to enable counterfactual policy evaluation despite demand censoring. In backlogging systems and lost-sales models with zero lead time, our algorithms adapt to unknown demand changes while matching, up to logarithmic factors, the rates known for the corresponding stationary learning problems. In lost-sales systems with positive lead times, the combination of demand censoring and delayed replenishment restricts counterfactual policy evaluation and leads to weaker regret guarantees. We complement the theoretical analysis with simulation results showing that our methods significantly outperform existing non-oracle benchmarks.
Data-Aware and Scalable Sensitivity Analysis for Decision Tree Ensembles
oai:arXiv.org:2602.07453v2
arXiv:2602.07453v2 Announce Type: replace-cross
Abstract: Decision tree ensembles are widely used in critical domains, making robustness and sensitivity analysis essential to their trustworthiness. We study the feature sensitivity problem, which asks whether an ensemble is sensitive to a specified subset of features -- such as protected attributes -- whose manipulation can alter model predictions. Existing approaches often yield examples of sensitivity that lie far from the training distribution, limiting their interpretability and practical value. We propose a data-aware sensitivity framework that constrains the sensitive examples to remain close to the dataset, thereby producing realistic and interpretable evidence of model weaknesses. To this end, we develop novel techniques for data-aware search using a combination of mixed-integer linear programming (MILP) and satisfiability modulo theories (SMT) encodings. Our contributions are fourfold. First, we strengthen the NP-hardness result for sensitivity verification, showing it holds even for trees of depth 1. Second, we develop MILP-optimizations that significantly speed up sensitivity verification for single ensembles and for the first time can also handle multiclass tree ensembles. Third, we introduce a data-aware framework generating realistic examples close to the training distribution. Finally, we conduct an extensive experimental evaluation on large tree ensembles, demonstrating scalability to ensembles with up to 800 trees of depth 8, achieving substantial improvements over the state of the art. This framework provides a practical foundation for analyzing the reliability and fairness of tree-based models in high-stakes applications.
Simultaneous estimation of multiple discrete unimodal distributions under stochastic order constraints
oai:arXiv.org:2603.11532v2
arXiv:2603.11532v2 Announce Type: replace-cross
Abstract: We study the problem of estimating multiple discrete unimodal distributions, motivated by search behavior analysis on a real-world platform. To incorporate prior knowledge of precedence relations among distributions, we impose stochastic order constraints and formulate the estimation task as a mixed-integer convex quadratic optimization problem. Experiments on both synthetic and real datasets show that the proposed method reduces the Jensen-Shannon divergence by 2.2% on average (up to 6.3%) when the sample size is small, while performing comparably to existing methods when sufficient data are available.
Stable GFlowNets with TV Monitoring and Probabilistic Guarantees
oai:arXiv.org:2605.01729v2
arXiv:2605.01729v2 Announce Type: replace-cross
Abstract: Generative Flow Networks (GFlowNets) learn to sample states proportional to an unnormalized reward. Despite their theoretical promise, practical training is often unstable, exhibiting severe loss spikes and mode collapse. To tackle this, we first assess the sensitivity of GFlowNet objectives, demonstrating that a small Total Variation (TV) distance between the learned and target distributions does not preclude unbounded training loss. Motivated by this mismatch, we establish converse guarantees by deriving loss-to-TV bounds that certify global fidelity from bounded trajectory balance losses. Lastly, we propose Stable GFlowNets, an algorithm that leverages our theoretical results to stabilize training, and empirically demonstrate improved training behavior and superior distributional fidelity.
Co-Design Optimization for Data Center Cooling System via Digital Twin
oai:arXiv.org:2605.15516v3
arXiv:2605.15516v3 Announce Type: replace-cross
Abstract: Liquid-cooled exascale supercomputers dissipate heat through cooling plants organized as multiple parallel subloops, but how to allocate coolant distribution units (CDUs) across subloops and how to distribute flow among them has not been systematically addressed for facilities at this scale. This paper presents a three-layer optimization framework that jointly determines the integer partition of CDUs across subloops, the continuous flow fraction allocation, and the per-timestep co-design optimization of total flow rate and supply temperature subject to per-subloop thermal safety constraints. The Modelica simulation model is built based on the data of the Frontier exascale supercomputer at Oak Ridge National Laboratory. By developing a reduced-order surrogate model, all 611 feasible partitions of 25 CDUs are evaluated across the full year operational dataset of 49,353 timesteps. Three progressively richer operational strategies are compared, ranging from flow control optimization to full three-layer co-design optimization with dynamically adjusted flow fractions. The optimal design within the surrogate optimization problem is a two-subloop plant achieving 35.48% annual cooling energy savings, only 0.18% above the current three-subloop design at 35.30%. Most of the savings are delivered by supervisory co-optimization of total flow rate and supply temperature; the distinct role of flow fraction optimization is design robustness rather than additional raw savings. Flow fraction optimization compensates for any feasible CDU-to-subloop assignment, reducing the design sensitivity by 93% and providing a low-cost software-only pathway to near-optimal performance on the existing Frontier hardware. The framework is transferable to other liquid-cooled high-performance computing plants.
An interpretable Good--Turing restart criterion for k-means++
oai:arXiv.org:2607.08243v2
arXiv:2607.08243v2 Announce Type: replace-cross
Abstract: The k-means++ algorithm is commonly restarted multiple times to avoid poor local optima, yet the number of restarts is almost always chosen arbitrarily and applied uniformly regardless of data set difficulty. This undermines any comparison relying on such a choice and wastes computation on easy data sets while potentially under-serving hard ones. Here, we introduce the Good-Turing Restart Criterion (GTRC). This combines a Good-Turing estimate, a proven unconditional bound, and a confidence-based bound on the probability that a further restart would improve on the current result, stopping once this probability falls below a user-specified tolerance. Our experiments on 34 real-world data sets show that GTRC identifies the point beyond which further k-means++ restarts yield only negligible improvement, achieving a more favourable balance between the number of restarts used and clustering quality than three existing stopping rules for multistart local search and popular fixed restart counts. Software: https://github.com/RCdeAmorim/Good-Turing-Restart-Criterion.
Optimizing the Preconditioner: A Black-box Online-to-Nonconvex Conversion with Static Regret Minimization Oracles
oai:arXiv.org:2607.17607v2
arXiv:2607.17607v2 Announce Type: replace-cross
Abstract: We study whether stochastic nonconvex optimization can be reduced to ordinary static regret minimization in online convex optimization in a black-box manner. For smooth nonconvex objectives, our reduction maintains a predictable gradient tracker, while a black-box online learner selects a preconditioner that determines how this tracker is transformed into the update direction. The learner receives linear convex losses and is evaluated against a single fixed comparator over one undiscounted online game. For a $\beta$-smooth objective with range bounded by $M$ and an unbiased stochastic-gradient oracle with variance bounded by \(\sigma^2\), we establish $$\frac{1}{T}\sum_{t=1}^T
\mathbb E\!\left[\|\nabla f(x_t)\|_2^2\right]
\lesssim
\frac{\sigma\sqrt{M\beta}}{\sqrt T}
+
\frac{\sqrt{M\beta}\,
\mathscr R_T(\mathcal A,I_d)}{T}
+
\frac{M\beta}{T}.$$ Consequently, any black-box OCO algorithm with $\mathscr R_T(\mathcal A,I_d)=O(\sqrt T)$ recovers the classical $O(\frac{1}{\sqrt{T}})$ convergence rate.
We further show that the same black-box framework extends beyond the smooth setting to Lipschitz nonconvex objectives without Lipschitz continuous gradients. Importantly, this extension continues to rely only on an ordinary static-regret guarantee and requires no stronger notion of online regret. When the OCO oracle admits square-root static regret, the resulting conversion achieves the optimal $O(T^{-2/7})$ convergence rate for the corresponding Goldstein stationary point. These results resolve the open problem posed by Chen and Hazan (2024). More broadly, our framework separates optimizer design into gradient prediction and online preconditioner selection, providing a principled perspective on how adaptive optimization methods may be understood through static regret and applied in nonconvex optimization.
Cautious optimism for deep parameterized quantum circuits
oai:arXiv.org:2607.21409v2
arXiv:2607.21409v2 Announce Type: replace-cross
Abstract: A central challenge in quantum machine learning is understanding the scaling behavior of parameterized quantum circuits (PQCs). In particular, it remains unclear how their performance on unseen data changes as the number of trainable parameters increases. Prior works have derived formal generalization guarantees for quantum models, but it is well-known that many such results do not fully characterize generalization behavior in practice. In this work, we show that gradient-based PQCs can exhibit improved performance on unseen data as model size increases, displaying the phenomenon of double descent. This contrasts with the traditional view that larger models lead to degraded generalization. We provide analytical results rigorously underpinning this behavior by leveraging add-one-in perturbation techniques and spectral properties of random matrices. We support these results with numerical experiments on re-uploading PQCs across several data sets and training set sizes, consistently observing the predicted double descent behavior. While other obstacles on the path toward practical quantum machine learning remain, our finding that deeper parameterized quantum circuits do not necessarily exhibit degraded performance provides reasons for cautious optimism.