• raseliarison
  • nirinA
  • adrien
  • blog
  • code
  • FAQ
  •  home  
  •  news  
    • arXiv
      • astro-ph
      • cond-mat
      • cs
      • eess
      • gr-qc
      • hep-ex
      • hep-lat
      • hep-ph
      • hep-th
      • math
      • math-ph
      • nlin
      • nucl-ex
      • nucl-th
      • physics
      • q-bio
      • quant-ph
      • stat
    • physics
      • phys.org
      • physics world
    • linux
      • kernel
      • slackware
    • nature
      • natcomputsci
      • natastron
      • natbiomedeng
      • nenergy
      • nnano
      • natmachintell
      • nbt
      • nmeth
      • natecolevol
      • nmicrobiol
      • ng
      • nchembio
      • natelectron
      • micronano
      • nphoton
    • bioRxiv
    • plos one
    • world
      • BBC
      • Al Jazeera
    • earth
      • earth observatory
      • weather
      • weather forecast
    • universe
      • apod
      • hubble
      • atel
      • nasa
  •  wiki  
  •  gemini  
  •  python  
  • stat updates on arXiv.org

    stat updates on the arXiv.org e-print archive.

    Extreme classification: beating chance with one training example from each class

    oai:arXiv.org:2609.20897v1

    arXiv:2609.20897v1 Announce Type: new Abstract: We study a minimal classification problem: Given independent labeled observations $X\sim P$ and $Z\sim Q$ from two unknown distributions $P,Q$, and given an independent target $Y$ drawn with equal probability from $P$ or $Q$, can one classify $Y$ strictly better than chance whenever $P\neq Q$? The one-nearest-neighbor rule succeeds for every pair of multivariate Gaussian distributions with distinct means and a common positive-definite covariance matrix but can perform strictly worse than chance even for smooth densities on the real line. We construct a fixed randomized kernel rule whose expected accuracy is exactly $1/2+\operatorname{MMD}_k^2(P,Q)/4$, and obtain characteristic kernels on countably generated measurable spaces from countable families of measurable binary questions. We also prove that a deterministic order rule on $\mathbb R$ beats chance for every pair of distinct Borel probability measures. A measurable encoding then gives a deterministic distribution-free rule which beats chance on every countably generated measurable space, in particular every separable metric space. Finally, we show that no rule works for every distinct pair of distributions and every unknown unbalanced class prior; under adaptive target-class selection, every rule other than a fair coin is strictly worse than chance for some finitely supported pair.

    https://arxiv.org/abs/2609.20897


    Identification problem and quasi-maximum likelihood estimation for matrix-variate CP-factor models

    oai:arXiv.org:2609.20910v1

    arXiv:2609.20910v1 Announce Type: new Abstract: Matrix-valued time series, arising in diverse fields such as economics, neuroscience, and recommender systems, have become increasingly prominent in modern data analysis. Among various modeling frameworks, the matrix-variate CP-factor model represents an important and widely applicable class for capturing low-rank structures in matrix time series. In this paper, we provide a unified treatment for the identification problem of CP factor models using both analytic and algebraic tools. In particular, we characterize the parameter space as the union of two subspaces, one identifiable and the other non-identifiable. We show that existing estimation methods only apply to an open proper subset of the identifiable subspace. In contrast, we propose a quasi-maximum likelihood estimation (QMLE) procedure for CP factor models, which allows consistent estimation on the whole identifiable subspace. Moreover, we show that the estimated loading matrices by QMLE achieve a faster convergence rate compared to existing approaches. A simulation study and a real application are conducted to demonstrate the finite-sample performance of the proposed method.

    https://arxiv.org/abs/2609.20910


    Complex Problem Solving in Large Language Models: A Statistical Control Survey and Diagnostic Framework

    oai:arXiv.org:2609.20973v1

    arXiv:2609.20973v1 Announce Type: new Abstract: Complex problem solving (CPS) with large language models (LLMs) is often framed as a matter of stronger reasoning or longer generation. Yet early-step error amplification, prompt brittleness, and failures to revise incorrect commitments are difficult to explain by missing knowledge or expressive capacity alone. This survey interprets CPS as a sequential estimation-and-decision problem over a latent solution state. A controller maintains a belief about an unobserved solution trajectory, updates it as noisy intermediate evidence arrives, and decides whether to commit, verify, branch, roll back, or abstain to minimize expected loss. Reasoning supplies candidate transitions and interpretations, whereas process control shapes and evaluates those proposals and regulates subsequent transitions and observations. Within this framework, we organize existing methods around five components: explicit state representation, transition structuring, validation and constraint enforcement, search and rollback, and uncertainty management. We also interpret evaluation metrics according to the statistical quantities they estimate. The framework further yields a diagnostic hypothesis: interventions should be most effective when they target the error or uncertainty component implicated by an observed failure. We distinguish systematic, stochastic, and irreducible error together with epistemic and aleatoric uncertainty, and call this alignment problem-control fit and its failure control mismatch. For example, additional sampling may reduce sampling variability while leaving a shared systematic error unchanged. This perspective clarifies what current methods estimate and control, what remains uncontrolled, and why reliable validation, targeted recovery, calibrated uncertainty, and matched-budget evaluation are central open problems.

    https://arxiv.org/abs/2609.20973


    Aggregated Posterior Predictive Checks for Generative Modeling

    oai:arXiv.org:2609.20999v1

    arXiv:2609.20999v1 Announce Type: new Abstract: Latent variable generative models are commonly fit using simple priors over latent variables, but draws from these priors often fail to produce realistic data. This failure is due to a mismatch between the prior and the aggregated posterior, the distribution of latent variables induced by the fitted model and the data. This mismatch is often viewed as evidence that the prior is misspecified and should be replaced. Alternatively, in modern generative models, a two-stage strategy is increasingly used where first, the model is fit, and second, the aggregated posterior is estimated (van den Oord et al.,2017; Rombach et al., 2022.). Synthetic data are then obtained by sampling from this aggregated posterior instead of the prior. To check such procedures, we introduce the aggregated posterior predictive check (APPC). Theoretically, we establish sufficient conditions under which the APPC is asymptotically calibrated. For probabilistic principal component analysis, we show that the APPC can remain calibrated under a misspecified latent prior when pervasive factors permit recovery of the signal space. Experiments with variational autoencoders show that aggregated posterior sampling improves generation for heavy-tailed and clustered data relative to Gaussian prior sampling while performing comparably to models with more flexible latent priors.

    https://arxiv.org/abs/2609.20999


    A Smoothed Discrepancy Principle for Random Feature Methods and Neural Networks

    oai:arXiv.org:2609.21017v1

    arXiv:2609.21017v1 Announce Type: new Abstract: We study data-driven early stopping for spectral regularisation methods in the classical non-parametric regression setting. Building on the discrepancy principle, we propose a multi-scale stopping rule that applies to general kernel estimators and show that, unlike previous approaches, it achieves full adaptivity over all smoothness levels in the well-specified case. A key contribution of our work is an extension based on random feature approximations, which reduces computational cost on large datasets while preserving minimax-optimal statistical guarantees. Our procedure not only selects an optimal stopping time but also provides a fully data-driven choice of the number of random features needed to achieve optimal rates. Through the established connection between random features and neural networks in the neural tangent kernel regime, our method further yields a principled, data-driven recommendation for the network width. We prove that the resulting simultaneously chosen width and stopping time allow neural networks to attain minimax-optimal learning rates without prior knowledge of smoothness or capacity parameters.

    https://arxiv.org/abs/2609.21017


    A New Generalized Birnbaum-Saunders Regression Model: Inference, Diagnostics, and Applications

    oai:arXiv.org:2609.21063v1

    arXiv:2609.21063v1 Announce Type: new Abstract: The Birnbaum-Saunders (BS) distribution has become a widely used model for positive and asymmetric continuous data. Several extensions of the BS distribution have been proposed, including regression formulations that relate its parameters to covariates. In this paper, we introduce a new regression model for the Generalized Birnbaum-Saunders (GBS) distribution, an extension that has received limited attention in the literature despite offering a stronger physical interpretation. The proposed framework follows the structure of Generalized Additive Models for Location, Scale, and Shape (GAMLSS), allowing covariate effects to be incorporated into all three parameters of the distribution, thereby providing greater flexibility for modeling data. In particular, one of the model parameters has a direct interpretation as the median of the response variable, facilitating the interpretation of covariate effects on the conditional median. Parameter estimation is performed via maximum likelihood, and hypothesis tests are proposed based on the Wald test. The proposed regression model is implemented in R through the gamlss package, allowing access to a broad range of tools for model fitting, diagnostics, and assessment. Monte Carlo simulation studies are conducted to investigate the finite-sample performance of the maximum likelihood estimators. The results show that the estimators exhibit desirable properties, with bias decreasing and efficiency increasing as sample size grows. Additionally, the behavior of the Wald test is investigated, showing good performance for moderate to large sample sizes. Finally, two applications to real data illustrate the practical usefulness of the proposed GBS regression model, demonstrating that it is a flexible alternative for modeling positive and asymmetric data compared to the BS model

    https://arxiv.org/abs/2609.21063


    Triply-Scalable Equivariant Gaussian Process Modeling

    oai:arXiv.org:2609.21085v1

    arXiv:2609.21085v1 Announce Type: new Abstract: Gaussian processes (GPs) provide principled probabilistic predictions while encoding prior knowledge, including equivariances. Yet, their use in large-scale scientific problems is limited by computational cost. Equivariant neural networks are common but typically lack the uncertainty quantification offered by GPs, which is valuable in applications such as molecular research. High-dimensional inputs and large symmetry groups further demand scalability. We establish results pertaining to the interplay of GP equivariance and conditioning and leverage them to obtain equivariant sparse GPs through suitable mean functions and covariance kernels. We instantiate this framework with a flexible class of integration-free equivariant kernels, yielding scalable and data-efficient GP inference. In particular, we introduce triply scalable equivariant Gaussian processes. We employ equivariant sparse variational Gaussian processes for $\mathrm{SO}(2)$-equivariant vector fields and molecular property prediction. Alongside the SVGP, we develop a matrix-free equivariant full-GP implementation that combines an exact Kronecker reduction with preconditioned conjugate-gradient solves, enabling fast and scalable evaluation of the full joint predictive density. We further compare different approaches for selecting inducing points in the equivariant sparse GP models. Our test cases include synthetic $\mathrm{SO}(2)$-equivariant fields as well as the prediction of electric dipole moments of N-methylformamide based on quantum chemistry simulations, achieving accurate, uncertainty-aware predictions at a fraction of the computational cost of classical GP inference.

    https://arxiv.org/abs/2609.21085


    Bayesian additive regression trees for evaluating treatment benefit predictors using observational data

    oai:arXiv.org:2609.21097v1

    arXiv:2609.21097v1 Announce Type: new Abstract: A treatment benefit predictor (TBP) is an algorithm that maps a patient's characteristics to their putative benefit from a given treatment, which can be used to inform treatment decisions. However, a TBP must be evaluated in the target population before being adopted for patient care. When only observational data are available, evaluating TBPs requires standard causal identification assumptions, as treatment assignment in such settings is not random. We obtain the posterior distributions of predictive performance measures to evaluate prespecified TBPs using observational data, by taking advantage of Bayesian additive regression trees (BART). We illustrate the evaluation of TBPs using selected measures and graphical visualizations: the concentration of benefit ($C_b$) index and the moderate calibration curve. Simulation studies of binary and continuous outcomes settings, including balanced and imbalanced treatment allocation in the binary setting, establish the validity of the proposed approach. In a case study, we use this approach to assess a TBP for systemic antibiotic therapy for patients with chronic obstructive pulmonary disease (COPD). We show that the constructed TBP does not in fact make calibrated predictions, because it both makes optimistic predictions of risks and exaggerates the risk reduction due to treatment. We conclude that flexible Bayesian approaches have the potential to assess TBPs, offering opportunities for flexible model specifications, adjustment for confounding, and uncertainty characterization.

    https://arxiv.org/abs/2609.21097


    Exact analysis of a split--merge queue with latent Erlang-factor dependent subtask times

    oai:arXiv.org:2609.21146v1

    arXiv:2609.21146v1 Announce Type: new Abstract: This paper studies a two-server split--merge queue with positively dependent subtask service times modeled through a latent-factor bivariate Erlang construction. An exact characterization of the split--merge completion time is obtained, including explicit formulas for its first two moments and the resulting mean waiting time. Under fixed marginal service-time distributions, independence is shown to stochastically increase the completion time and hence overestimate mean waiting time. Numerical illustrations show that this benchmark gap can be substantial.

    https://arxiv.org/abs/2609.21146


    Repulsive normalizing flow mixtures for adaptive importance sampling: reliability analysis of complex systems

    oai:arXiv.org:2609.21160v1

    arXiv:2609.21160v1 Announce Type: new Abstract: Accurate rare-event estimation can be computationally expensive. Classical adaptive importance sampling (IS) schemes often rely on restrictive proposal families and can struggle under multiple failure modes. We propose FAMIS, a flow-based multiple importance sampling (MIS) framework that learns a nonuniform mixture of normalizing flow proposals for rare event estimation. The method does not require presampled failure data or prior knowledge of the number, location, or geometry of the failure modes. Instead, it adaptively learns the mixture through sequential evaluations of the limit state function. To guide training toward the failure domain, FAMIS uses a smooth rare-event surrogate and a tempered target sequence. A defensive exploration mixture improves early-stage coverage, a Rao Blackwellized update adapts the mixture weights, and a Jensen-Shannon repulsion term promotes separation and diversity among the base components. The final failure probability is computed with a deterministic-mixture MIS estimator. Numerical experiments demonstrate that FAMIS accurately approximates quasi-optimal IS densities with fewer training samples and model evaluations, providing stable variance reduction across complex reliability problems.

    https://arxiv.org/abs/2609.21160


    Where quantum and quantum-like methods earn their place in behavioural-trial analysis: three simulation pilots across the RCT pipeline

    oai:arXiv.org:2609.21269v1

    arXiv:2609.21269v1 Announce Type: new Abstract: Behavioural randomized controlled trials (RCTs) are the empirical workhorse of behavioural science, yet quantum computing has reached the field only through opinion-dynamics demonstrations and decision-theory formalism, leaving the RCT analysis pipeline unexamined. I ask, stage by stage, where quantum or quantum-like computation earns its place against a strong classical baseline, and report three controlled simulation pilots across the pipeline. In a power-calculation pilot anchored on a 97-outcome preprocessing multiverse of real behavioural RCTs, quantum amplitude estimation recovers enumerated multiverse-significance fractions with absolute error 20-30 times smaller than classical sampled Monte Carlo at matched query budgets on the two non-degenerate outcomes, conditional on efficient state preparation. In an interference pilot, an entanglement-structured estimator recovers a four-way spillover that a pairwise model cannot represent at any sample size. In an adaptation pilot, a quantum-inspired contextual bandit is decisively beaten by a correctly specified classical baseline. The three results - one conditional win, one representational win, one honest loss - yield a decision rule for when quantum methods are worth adopting, piloting, or deferring, together with a five-gap research agenda.

    https://arxiv.org/abs/2609.21269


    The Manifold Hypothesis under Unknown Gaussian Noise:Conditional Certificates and Consistent Dimension Estimation

    oai:arXiv.org:2609.21311v1

    arXiv:2609.21311v1 Announce Type: new Abstract: We study what noisy data can establish about the Manifold Hypothesis under explicit identification and regularity conditions. A population residual certificate combines independent-view localization, Gaussian concentration, membership uncertainty, and population transfer. Existing rectifiability criteria then yield a covered-scale consequence. For a local smooth manifold with positive H\"older density, the actual-ball covariance limit identifies the spectral crossing with geometric dimension. We prove almost-sure eventual recovery under repeated observations. Reusing accurate localization averages improves the sufficient point-sample condition from $Nr^{d+4}\gg\log N$ to $Nr^d\gg\log N$, with replication $kr^2\gg\log N$. A two-mass certificate controls incorrect geometric-dimension emissions under declared class bounds. For single observations with unknown Gaussian noise, affine-support or known coordinate-bound restrictions provide noise intervals and consistent Gaussian correlation-dimension estimators. Ahlfors regularity identifies this exponent with Hausdorff dimension and with the geometric dimension of a homogeneous smooth class. Exact Cantor calculations delineate the limits of integer spectral counts and adjacent-radius slopes. We credit established local PCA, rectifiability, concentration, binomial inference, and deconvolution results before specifying our constructions. Reproducible experiments distinguish point estimation, finite-scale coverage, and certificate emission.

    https://arxiv.org/abs/2609.21311


    Diagonalized Attention for Individualized Regression: Latent-Row Localization and Prediction

    oai:arXiv.org:2609.21320v1

    arXiv:2609.21320v1 Announce Type: new Abstract: Modern text and image representations are often matrix-valued, with rows corresponding to tokens, patches, or other local feature vectors. Predictive information is often sparse but sample-specific, making classical sparse regression methods with a common support poorly suited to this heterogeneity. This paper formalizes an individualized sparse regression framework for matrix-valued covariates in which each observation has its own rows of interest, while the associated regression effects are shared across the population. To estimate this model, we introduce a diagonalized attention mechanism that uses query--key scores to localize sample-specific signal rows and a value matrix for downstream regression. The proposed method has a parameter dimension independent of sample size and can identify rows of interest for new observations without their responses. We establish existence theorems showing that, under suitable score-separation and concentration conditions, single-head and multi-head diagonalized attention models recover the latent rows with high probability, yielding prediction risk bounds. Our theory therefore provides a statistical explanation of how attention-based scoring localizes sample-specific signals in heterogeneous matrix-valued data. Simulations demonstrate strong prediction and localization in regression and misspecified classification across varying sample sizes, dimensions, and signal cardinalities. Real sentiment analyses show improved classification accuracy and interpretable token selection.

    https://arxiv.org/abs/2609.21320


    Sparse Identification for Automatic Large-Scale Screening: A Constraint-Aware Framework with Ultra Fast Decoding Algorithm

    oai:arXiv.org:2609.21321v1

    arXiv:2609.21321v1 Announce Type: new Abstract: In the early stages of a pandemic, identification of a small number of infected individuals through large-scale screening is critical for pandemic control, yet remains challenging under limited reagents and testing capacity. Existing group testing methods suffer from either high computational complexity or low identification accuracy. Even worse, no available methods provide theoretically rigorous analysis for sparse identification with hard constraints caused by the sample usage constraint and the dilution effect existing ubiquitously in practical applications. In this article, we propose the Logic Screening method (LoSc), an ultra fast, accurate, and theoretically grounded framework for large-scale screening. LoSc introduces a novel decoding algorithm with a very simple selection strategy, achieving identification of all positives with only O(klogn) pooled tests. The decoding relies only on logical operations, enabling direct hardware implementation and yielding ultra fast computational implementation. Moreover, LoSc explicitly incorporates dilution and sample usage constraints into pooling designs, and establishes theoretical guarantees to guide optimal pooling configurations. Extensive simulations confirm the superior effectiveness, efficiency, and scalability. We believe LoSc offers a fast and reliable solution for automatic large-scale screening.

    https://arxiv.org/abs/2609.21321


    Asymptotic Anytime-Valid Quantile Inference under Local Differential Privacy

    oai:arXiv.org:2609.21338v1

    arXiv:2609.21338v1 Announce Type: new Abstract: Sequential quantile inference is difficult under local differential privacy because every record is randomized before reaching the analyst and the limiting quantile variance depends on an unknown density. We develop an online procedure that combines randomized response with dynamically chained parallel stochastic gradient descent (P-SGD). The resulting Polyak--Ruppert estimator admits a strong Gaussian approximation. A cross-chain quadratic statistic, computed entirely from private iterates, consistently estimates the limiting variance without a separate online density estimator. These results yield asymptotic confidence sequences and, under polynomial chain growth, asymptotic time-uniform coverage. Arm-wise constructions support locally private quantile best-arm identification, time-uniform simple-regret bounds, and sequential A/B tests of quantile treatment effects. Simulations and salary-data analyses illustrate the finite-sample behavior and practical use of the proposed methods.

    https://arxiv.org/abs/2609.21338


    Robust Dual-Regularized Variable Selection under Outlier Contamination

    oai:arXiv.org:2609.21342v1

    arXiv:2609.21342v1 Announce Type: new Abstract: Real data often contain unusual observations that can exert disproportionate effects on variable selection, especially in complex predictor settings. We propose a two-stage {\it sparse median outer product of gradients (smOPG)} method for variable selection in single index models with outlier contamination. We first estimate sparse local gradients via \(\ell_1\)-penalized local median regression and then recover the active predictor set from a rank-one sparse approximation of the resulting gradient matrix using regularized singular value decomposition. The combination of median regression and local weighting provides robustness to both response outliers and leverage points. Extensive simulations across varying dimensions and contamination mechanisms demonstrate the favorable variable selection performance of smOPG relative to existing methods. Applications to air pollution and genomic data demonstrate practical utility, while theory establishes active-set recovery without requiring selection consistency of individual local regressions.

    https://arxiv.org/abs/2609.21342


    Brownian Heads for Deep ReLU Representations: Activation Mass and the Cost of Same-Sample Selection

    oai:arXiv.org:2609.21422v1

    arXiv:2609.21422v1 Announce Type: new Abstract: Deep representation learning often selects hidden features and fits the final predictor on the same sample, so fixed-feature analysis performed after selection can omit selection cost. We study the conditional empirical Rademacher complexity of deep ReLU representations followed by bounded-norm predictors in additive or L\'evy-Brownian RKHSs, termed Brownian heads. For a fixed representation, we derive an exact dual identity and sharp bounds in terms of activation mass, the average norm of the observed hidden vectors. Under same-sample selection, the representation supremum induces a quadratic Rademacher process. Brownian layer-cake and Gaussian-projection identities reduce it to coordinatewise or signed projected threshold traces, separating realized scale from selection complexity. For samples with pairwise-distinct inputs, explicit scalar ReLU families match the finite-trace and VC rates up to universal constants at the realized trace-and-envelope level. Induced-norm contraction also yields architecture-level bounds for rectangular, rank-deficient ReLU networks. Experiments verify the sharp bounds and rates, exhibit a selection gap at fixed activation mass, and assess the predictive feasibility of Brownian heads.

    https://arxiv.org/abs/2609.21422


    Robust tests for log-normal lifetimes under cyclic-stress accelerated tests and interval monitoring

    oai:arXiv.org:2609.21434v1

    arXiv:2609.21434v1 Announce Type: new Abstract: Some products are, by nature, designed to operate under stress conditions that repeatedly cycle between two levels, such as batteries undergoing repeated charge and discharge, or components exposed to recurring thermal or pressure cycles. When such products are tested under accelerated conditions, the applied stress is naturally elevated in the same cyclic manner rather than held constant or increased in a single step. Cyclic-stress accelerated life testing (CyALT) reflects this by alternating the applied stress between specified levels. Units are often inspected only at fixed times, so the data consist of interval failure counts rather than exact lifetimes. Hypothesis testing plays a crucial role in this setting. It can assess whether the applied stress significantly affects lifetime, or validate the parameter values used to design the experiment. Classical tests rely on the maximum likelihood estimator, which is highly sensitive to contamination. Since interval probabilities are estimated jointly across stress conditions, a single anomalous failure count can distort the entire parameter vector and severely affect the test's significance level and power. This paper develops Wald-type and Rao-type test statistics for log-normal CyALT lifetimes under interval monitoring, based on the weighted minimum density power divergence estimator. Both tests are derived for simple and composite null hypotheses. A Monte Carlo study shows that classical tests lose control of the significance level rapidly under contamination, while the robust tests stay close to nominal with only a modest loss of power under clean data. An application to air-conditioner evaporator lifetime data further illustrates the proposed tests.

    https://arxiv.org/abs/2609.21434


    Improving the Predictive Performance of Bootstrap Aggregating by Dirichlet Resampling

    oai:arXiv.org:2609.21454v1

    arXiv:2609.21454v1 Announce Type: new Abstract: We revisit Breiman's observation that reducing inter-tree correlation without weakening individual trees can improve random forests. Building on this principle, we introduce two variants: Dirichlet-Multinomial Bagging Random Forest (DM) and Dirichlet-Weighted Random Forest (DW). Both modulate sample reweighting via a concentration parameter $\alpha>0$. We provide a simple theoretical criterion that clarifies when these variants behave indistinguishably from standard random forests, and we use it to guide a lightweight tuning strategy. In a controlled evaluation on public classification benchmarks, DM and DW are consistently competitive and often stronger than other random-forest (RF) baselines, with negligible additional runtime.

    https://arxiv.org/abs/2609.21454


    Scentree: a framework for generating scenario trees for multistage stochastic programming

    oai:arXiv.org:2609.21495v1

    arXiv:2609.21495v1 Announce Type: new Abstract: We present scentree, an open-source Python package for constructing a scenario fan and a scenario tree for multistage stochastic programming from historical data. It combines machine learning and multivariate time series models to obtain a scenario fan that captures inter-stage dependencies in the stochastic processes. This scenario fan is subsequently transformed into a scenario tree suitable for multistage stochastic optimization, providing a flexible and extensible framework for uncertainty modeling. A key contribution is the automation of the complete workflow, including model selection, parameter estimation, scenario fan generation, and scenario tree construction. Scentree does not rely on assumptions about the underlying data distribution, reducing the statistical expertise required to produce a scenario tree. Furthermore, it is agnostic to the specific multistage stochastic problem to be solved.

    https://arxiv.org/abs/2609.21495


    Estimating heterogeneous treatment effects from randomised trials: a comparison of the risk modelling and treatment effect modelling approaches

    oai:arXiv.org:2609.21526v1

    arXiv:2609.21526v1 Announce Type: new Abstract: Introduction Risk modelling (RM) and treatment effect modelling (EM) are two approaches to build models to estimate heterogeneous treatment effects from randomized control trials (RCT). RM is a two-stage approach; it estimates a predicted baseline risk at the first stage and then includes it as the only effect modifier in a regression model in the second stage. EM is a full treatment interaction multivariable model. Both approaches have theoretical advantages and limitations, but a thorough comparison including a simulation study is missing. Methods In a theoretical review of the two approaches, we present their underlying assumptions and theoretical advantages and disadvantages. Based on our theoretical review, we design simulation scenarios for RCTs with a dichotomous outcome and evaluate the performance of both approaches with respect to the root mean square error and the bias in the predicted risk difference. Results In the theoretical part we argue that baseline risk is a treatment effect modifier in many clinical situations. RM is a dimensionality reduction approach which, however, makes strong assumptions about the role of prognostic factors modifying the treatment effect. EM's greater flexibility is a possible advantage when the sample size is large. In most simulated scenarios RM performs better than EM, even when the assumptions underlying RM are not fully met. The advantage of RM diminishes as sample size increases. Conclusion When choosing between RM and EM the available sample size and the plausibility of their underlying assumptions should be considered.

    https://arxiv.org/abs/2609.21526


    Fuzzy entropic triple k-means

    oai:arXiv.org:2609.21538v1

    arXiv:2609.21538v1 Announce Type: new Abstract: This paper proposes fuzzy entropic triple k-means (FE3KM), an entropy-regularized fuzzy partitioning method for three-way data arrays that simultaneously clusters objects, variables, and occasions. Unlike standard fuzzy clustering, which controls fuzziness through a nonlinear exponent m, FE3KM embeds memberships linearly within a Least-Squares (LS) objective, using row-stochastic membership matrices and entropy regularization to represent uncertainty in cluster assignments. This linear structure is what enables an exact decomposition of total deviance into within- and between-cluster components, further attributable to each mode (objects, variables, occasions) and to individual clusters within each mode. This is an interpretive property unavailable to exponent-m fuzzy models. We derive Alternating Least-Squares algorithms for FE3KM, including an accelerated variant for large arrays, and formally relate the method to tandem clustering procedures and to Tucker3/three-mode partitioning models. A simulation study assesses recovery accuracy, robustness to noise and fuzziness mis-specification, and performance against tandem and three-way clustering baselines. An application to a benchmark three-way TV ratings dataset demonstrates how FE3KM recovers interpretable object-variable-occasion structures and quantifies the contribution of each cluster and mode to overall data variability.

    https://arxiv.org/abs/2609.21538


    Model-based estimation and imputation with torus missing values

    oai:arXiv.org:2609.21549v1

    arXiv:2609.21549v1 Announce Type: new Abstract: This paper addresses the problem of parameter estimation and model-based imputation for multivariate circular data lying on a p-dimensional torus in the presence of missing values. Actually, the periodic nature of the sample space invalidates conventional imputation techniques designed for Euclidean data. Then, we propose a general framework for maximum likelihood estimation under the wrapped elliptically symmetric family of distributions, with particular interest in the multivariate wrapped normal distribution, when the missing data mechanism is ignorable. The methodology leverages the conditional properties of the elliptically symmetric distributions on the unwrapped space, embedding the imputation of missing torus data into an Expectation-Maximization algorithm that treats both the wrapping coefficients and the missing entries as latent variables. Derivation of both the E and M steps is detailed and a working algorithm is discussed. Imputation methods are also taken into account. The finite-sample performance of the maximum likelihood estimator under ignorable missingness is assessed through Monte Carlo simulations under the wrapped normal specification. The methodology is also illustrated on data with artificially introduced missingness.

    https://arxiv.org/abs/2609.21549


    Phasing out Monte Carlo: exact acceptance probabilities for mean-spread release rules

    oai:arXiv.org:2609.21577v1

    arXiv:2609.21577v1 Announce Type: new Abstract: A large family of regulated release decisions accepts a production lot if and only if the pair (sample mean, sample standard deviation) of $n$ units falls in a fixed plane region. Content uniformity (USP <905>), capability release, variables sampling and percent-within-limits highway acceptance all take this form. The acceptance probability under a hypothesized unit-level law is not, so far as we are aware, evaluated exactly in current practice: for a general parent the joint finite-$n$ law of $(\bar X,s)$ is a constrained $n$-fold integral (Craig 1932 at $n=3,4$; Springer 1953 at general $n$) which admits no elementary reduction. Yet that probability is fixed by the law of the additive pair $(T_1,T_2)=(\sum_{i=1}^{n} X_i,\sum_{i=1}^{n} X_i^2)$, whose characteristic function is the $n$-th power of a one-observation transform. One two-dimensional Fourier inversion therefore delivers it for an arbitrary smooth parent - exactly, as an identity, and numerically on a grid whose dimension stays two whatever $n$ is - a deterministic alternative to both the simulation and the Normality assumption. The identity is exact; what we evaluate is a finite-grid inversion of it, whose error we report per example rather than bound. At $n=3$ the Uniform parent yields the joint law of $(\bar X,s)$, the acceptance probability and the capability law in closed form; we derive these, use them as the anchor of a validation ladder, and then work three rules end to end under non-normal parents at their prescribed sample sizes ($n=10$-$30$). Against the exact answer the Normal-theory incumbent errs by enough to change the risk a plan is designed against, and errs in both directions along the operating curve, so no single constant corrects it. In general only the full distribution suffices.

    https://arxiv.org/abs/2609.21577


    Testing Conditional Stochastic Dominance via Copula Derivatives

    oai:arXiv.org:2609.21622v1

    arXiv:2609.21622v1 Announce Type: new Abstract: Comparing two populations at the same physical covariate value requires more than conditional means or isolated target-point decisions: researchers may need evidence about an entire conditional-distribution ordering over a continuum, even when covariate margins differ. This paper makes that common-value comparison estimable under an explicit structure--flexibility tradeoff and turns the resulting surface into simultaneous evidence for first-order stochastic dominance. Population-specific margins map the common covariate value into each group, while a fitted copula-derivative representation links conditional distributions across the region. Uniform inference propagates uncertainty from both the margins and dependence model through a one-sided statistic with unknown binding locations. Under correct specification within a finite copula class, smoothness and trimming conditions, and a uniquely best candidate family, the procedure admits uniform control and consistent calibration. Simulations show increasing rejection as alternatives become more distinguishable, alongside model-selection sensitivity and small-sample size distortion. In a descriptive PSID application, the high--low parental-education comparison satisfies the two-direction criterion after multiplicity adjustment, whereas adjacent education-group comparisons remain inconclusive. The framework therefore supports region-wide distributional comparison while making its structural and inferential boundaries explicit.

    https://arxiv.org/abs/2609.21622


    A Partial Fay-Herriot model for Small Area Estimation: Estimating district-level consumption in Mozambique

    oai:arXiv.org:2609.21701v1

    arXiv:2609.21701v1 Announce Type: new Abstract: This paper proposes a new small area estimation approach that integrates Partial Least Squares within the Fay-Herriot model to address the challenges posed by high-dimensional and highly correlated auxiliary variables sets. The resulting Partial Fay-Herriot (PFH) estimator constructs supervised components that maximize their association with the target variable, enhancing model stability and predictive efficiency. Monte Carlo simulations demonstrate that PFH estimator achieves lower mean squared error than the standard Fay-Herriot estimator and outperforms principal components-based alternatives while relying on fewer latent dimensions. The methodology is applied to the estimation of district-level per capita consumption in Mozambique, where the survey data source is complemented by large set of correlated census variables. The resulting estimates highlight pronounced geographic heterogeneity and uncover spatial clusters of deprivation. Overall, the findings show that the proposed supervised dimension-reduction approach represents an effective and easily interpretable tool for producing reliable indicators in high-dimensional contexts.

    https://arxiv.org/abs/2609.21701


    Parameter estimation for graphon-interacting particle systems from discrete observations

    oai:arXiv.org:2609.21710v1

    arXiv:2609.21710v1 Announce Type: new Abstract: In this paper, we address the joint parameter estimation of drift and diffusion coefficients for heterogeneously interacting particle systems. Unlike the homogeneous setting, the interactions are governed by a graphon-weighted mean-field framework, which introduces significant analytical complexity: while homogeneous systems yield i.i.d. limits, our setting results in particles that are independent but non-identically distributed in the limit. Based on discrete observations of the system over a fixed time interval $[0, T]$, we propose a contrast function based on a pseudo-likelihood approach. We prove the consistency of the resulting estimators as the discretization step $\Delta_n \to 0$ and the number of particles $N \to \infty$. Furthermore, we establish asymptotic normality under the additional constraint $N \Delta_n \to 0$, demonstrating that the underlying graphon structure can be rigorously handled despite the lack of identical distribution in the limit.

    https://arxiv.org/abs/2609.21710


    Neural composite likelihood estimation: simulation based inference for time series

    oai:arXiv.org:2609.21762v1

    arXiv:2609.21762v1 Announce Type: new Abstract: Simulation based inference (SBI) circumvents the challenge of intractable likelihoods by using a simulator that generates data given parameter values. For instance, neural likelihood estimation (NLE) estimates the likelihood function by training a neural network to perform conditional density estimation on simulated data given corresponding parameters. However such density estimation is only feasible for relatively low dimensional data. We extend the scalability of SBI methods to a higher dimensional problem: long sequences with a complex dependency structure. We introduce Neural Composite Likelihood Estimation (NCLE). This divides the sequence into smaller, equal-sized batches. Instead of training NLE to estimate the likelihood for an entire sequence, we estimate the likelihood for each batch separately. The product of these forms an approximate composite likelihood (CL), and we perform frequentist inference using methods from the CL literature: we get a point estimate from maximising the approximate CL and obtain confidence intervals by estimating the Godambe information matrix. We demonstrate the effectiveness of NCLE with experiments on time series models.

    https://arxiv.org/abs/2609.21762


    How Many Posterior Samples? Calibrated Stopping for Adaptive Sensing

    oai:arXiv.org:2609.21813v1

    arXiv:2609.21813v1 Announce Type: new Abstract: In classification-oriented adaptive sensing, posterior samples characterize uncertainty at the current measurement state and can serve two roles: they may guide the next sensing direction, while their class labels provide votes for the candidate classes and determine whether sensing should continue. We focus on the stopping layer that turns these votes into a declaration, without modifying the posterior sampler or sensing directions. A natural plug-in rule declares when the observed vote share exceeds a threshold. We show that this threshold is not itself a confidence guarantee: when the underlying vote mass equals the threshold, the plug-in rule declares about half the time. As alternatives, we calibrate a fixed-sample rule and a finite-horizon sequential rule to a prescribed false-declaration probability, and study exact curtailment, which stops a fixed-pool rule once its final verdict is forced. We then derive how one-round declaration probabilities determine posterior-sample cost and classification accuracy along a sensing path. On MNIST with DDRM and a fixed PCA-guided probe sequence, curtailment saves up to 62% of posterior samples. Among the evaluated rules at matched operating points, sequential stopping reduces the cost the most. At a high accuracy, that same sequential rule can trade more posterior samples for fewer measurements.

    https://arxiv.org/abs/2609.21813


    Optimal study interior partitioning (OSIP): An optimization approach to observational study design under fine balance constraints

    oai:arXiv.org:2609.21902v1

    arXiv:2609.21902v1 Announce Type: new Abstract: Covariate matching between treatment and control groups remains a foundational methodology in observational causal inference, designed to replicate the covariate balance of randomized controlled trials. Despite substantial methodological advancements, applied researchers continue to face an inherent trade-off between achieving rigorous balance and preserving the statistical power of the matched sample. We introduce a novel framework for a matching method named Optimal Study Interior Partitioning (OSIP). OSIP frames observational matching as a joint optimization problem governed simultaneously by a maximum cardinality constraint on the number of strata, and a strict fine balance threshold on the matching carried out in each strata. OSIP is based on two-stage optimization algorithm. In both stages it considers a projection of the (high-dimensional) data into one-dimensional space by evaluating an estimated propensity score for each unit. The method is based on enforcing a partition of the input to subsets of consecutive intervals of estimated propensity scores. The novelty arises because we consider this task as a global optimization problem. In the first stage, we employ an optimal partitioning algorithm with respect to an auxiliary (one-dimensional) goal function. In the second phase, we go back to the original distance function and apply heuristics approaches locally optimizing the solution from the first stage. We compare our framework with established methodologies from the literature across well-studied empirical benchmarks. We use the standard benchmark of Lindner cardiovascular dataset to demonstrate that OSIP provides a highly efficient sweet spot solution for the corresponding multi-criteria goal. This conclusion is verified by considering also other benchmark datasets.

    https://arxiv.org/abs/2609.21902


    Riemannian Simultaneous Inference for Tangent Vector Field Regression

    oai:arXiv.org:2609.21910v1

    arXiv:2609.21910v1 Announce Type: new Abstract: We consider nonparametric tangent vector field regression on a Riemannian manifold without boundary. Because responses at different points lie in different tangent spaces, the proposed kernel estimator first parallel transports nearby responses to the target tangent space and then forms a volume-corrected local average. We first derive its uniform second-order bias, finite-bandwidth covariance, and stochastic rate. For simultaneous inference, the tangent norm is written as a supremum over the unit tangent bundle. Exact covariance whitening gives a unit-variance Gaussian field whose correlation length is of order $h$ along the base manifold and of order one along the fibre. Its local covariance geometry leads to a Gumbel limit with an explicit intrinsic constant. Combining this limit with Gaussian approximation and cross-fitted covariance estimation yields a feasible simultaneous confidence tube for the regression field. We further discuss improved finite-sample inference with bandwidth selection and high-order bias corrections. Simulations on various manifolds support the proposed inference procedure. A randomized reconstruction of global wind data illustrates how the tube's cross-sections describe spatially varying uncertainty.

    https://arxiv.org/abs/2609.21910


    A Probabilistic Modeling Framework for Transient Debris Outcomes in Low Lunar Orbits

    oai:arXiv.org:2609.21955v1

    arXiv:2609.21955v1 Announce Type: new Abstract: This work presents a novel probabilistic framework for the assessment of orbital debris fragment outcomes in the low lunar orbit (LLO) regime over time horizons ranging from a few hours to a hundred days after a debris generation event. The framework develops continuous distributions for the probability of sinking (lunar collision) and non-sinking over these transient time horizons, building upon NASA's Standard Breakup Model to provide insights into the variations in the likelihood of debris outcomes across the LLO regime and offering an alternative to computationally expensive Monte Carlo simulations. The effects of perturbative forces such as solar radiation pressure are used to assess the realization of debris outcomes over varying time horizons, providing an analytical framework that bounds the likelihood of each outcome. Results indicate variations in the probability of sinking over short time horizons depending on the originating location of the debris generation event, as well as a link between the physical characteristics of fragments and their likelihood of sinking over longer time horizons. The future incorporation of this framework into broader-scale orbital environment models, mission risk assessment procedures, and policy development are briefly discussed.

    https://arxiv.org/abs/2609.21955


    Schedule optimization for tau-leaping in masked discrete diffusion

    oai:arXiv.org:2609.21960v1

    arXiv:2609.21960v1 Announce Type: new Abstract: Masked discrete diffusion models are commonly accelerated using the so-called tau-leaping discretization method, which reveals several coordinates in parallel at each sampling step. The sampler replaces the joint conditional law of each revealed block by a product distribution, incurring a factorization error $\varepsilon_\text{fact}$ present even with perfectly learned predictors. We analyze the standard sampler on $N$ coordinates with $K$ sampling steps, whose random block sizes depend on a denoising schedule. Our analysis uses an exact integral representation of $\varepsilon_\text{fact}$ in terms of a distribution-dependent dependence density $\rho$, which records how conditional dependence evolves as the revealed fraction of coordinates grows. We develop estimators for this profile and quantify how estimation errors affect schedule selection. We derive recursive stationarity equations for the finite-$K$ optimization problem and, under a monotonicity condition, characterize its unique optimizer. In the joint limit $N,K\to\infty$, we obtain an explicit characterization of the optimal limiting smooth schedule and quantify the cost of random block sizes relative to a deterministic planner. When $\rho_N$ converges uniformly to a strictly positive continuous profile, optimizing over fixed smooth schedules can improve the leading constant but not the $N/K$ scaling of $\varepsilon_\text{fact}$. By contrast, if $\rho_N$ degenerates, suitable schedules can improve the asymptotic order relative to the uniform schedule. Examples based on stationary processes and exchangeable mixtures illustrate these two regimes.

    https://arxiv.org/abs/2609.21960


    Consistency of Optimal Matching-based Clustering for Mixtures of Markov Chains

    oai:arXiv.org:2609.21985v1

    arXiv:2609.21985v1 Announce Type: new Abstract: We study clustering of categorical sequences using the Optimal Matching (OM) distance under finite mixtures of finite-state Markov chains. We show that the normalized OM distance between two independent chains converges almost surely to a deterministic population quantity, concentrates exponentially around its finite-horizon mean, and admits an $O(\sqrt{\log n/n})$ convergence rate when the two chains have the same transition kernel. These population quantities yield a natural separation condition: the largest within-component limit must be smaller than the smallest between-component limit. Under this condition, hierarchical clustering with any bracketed linkage and Partitioning Around Medoids consistently recover the latent mixture partition. We also propose a consistent estimator of the number of components based on empirical OM distance profiles. The results extend to finite-state hidden Markov models and multichannel categorical observations. Overall, they provide a statistical justification for standard OM-based clustering methods for categorical time series.

    https://arxiv.org/abs/2609.21985


    Integrating Multivariate Adaptive Regression Splines into Small Area Estimation for Nonlinear Poverty Modeling in Java Island

    oai:arXiv.org:2609.21987v1

    arXiv:2609.21987v1 Announce Type: new Abstract: Small area estimation (SAE) modeling is often used to address the unreliability of estimates in areas with limited survey samples. However, standard SAE models (Fay-Herriot) can only accommodate linear patterns, whereas in real cases, including poverty, many patterns are nonlinear. This study proposes the implementation of a supervised learning model, multivariate adaptive regression splines (MARS), integrated within the SAE framework (MARS-SAE) to address the challenge of unreliable direct estimation, particularly in nonlinear cases. The hyperparameter tuning results yielded the best model with the following hyperparameter combinations: maximum basis function (Max BF) = 10, maximum interaction (MI) = 2, minimum observation (MO) = 1, penalty = 3, with general cross validation (GCV) = 6.128. MARS-SAE demonstrated very strong performance; empirically, the model proved more efficient at reducing estimation error due to its higher relative efficiency (RE) compared to the SAE Fay-Herriot model, as well as better reduction of relative standard error (RSE) than the comparison model, and it yielded a lower relative root mean squared error (RRMSE) compared to the SAE Fay-Herriot model (13.64% vs. 13.75%). MARS-SAE has been proven to capture nonlinear patterns very well while still maintaining interpretability, making it an effective tool for the predictive and diagnostic analyses of phenomena. Thus, this research can serve as an important reference for readers regionally and in Southeast Asia, given that Java Island is a strategic region for Indonesia, the largest economic power in the ASEAN region.

    https://arxiv.org/abs/2609.21987


    A Novel Convolution-Based Stratified Attribute Estimator for QRE Determination in R&D Tax Credit Studies

    oai:arXiv.org:2609.21998v1

    arXiv:2609.21998v1 Announce Type: new Abstract: The One Big, Beautiful Bill Act reinstates immediate expensing of domestic research expenses, reducing the tax burden and incentivizing reinvestment in U.S.-based research and development including software development across a wide range of industries. In this context, accurate and defensible estimation of qualified research expenses (QREs) is of high importance. Statistical sampling provides a practical framework for estimating QREs for a well-defined population of business components (sampling frame) documented in accordance with Internal Revenue Code Section 41 (Form 6765). IRS Revenue Procedure 2011-42 permits both attribute and variable statistical methods, although the latter (stratified mean and difference estimators) are often regarded as the standard approach to QRE estimation. In this paper, we compare attribute and variable statistical methods within a simulation study and demonstrate that attribute methods offer several theoretical and practical advantages. A key concern with the simple attribute estimator is the possibility of upward bias in QRE determination arising from a preponderance of high potential QRE (pQRE), non-qualified projects in the sampling frame. We address this concern by introducing a stratified sampling design that ensures adequate representation of high pQRE projects in the sample. We demonstrate that a convolution-based stratified attribute estimator produces a valid one-sided 95% lower confidence bound on total QREs across a range of sampling frame structures.

    https://arxiv.org/abs/2609.21998


    RKHS-Based Inference for Nonlinear Granger Causality via Conditional Centering

    oai:arXiv.org:2609.22007v1

    arXiv:2609.22007v1 Announce Type: new Abstract: Granger causality is commonly formulated through linear prediction in vector autoregressive models, limiting its ability to detect nonlinear predictive relationships. We propose a reproducing kernel Hilbert space (RKHS)-based test for nonlinear Granger non-causality in conditional mean for nonlinear autoregressive processes. The key idea is a conditional-centering decomposition of the target regression function into an own-history component and an orthogonal component capturing the additional predictive contribution of the potential source history. Non-causality is characterized by the vanishing of the latter component. The own-history component is estimated by kernel ridge regression, and the residuals are embedded in a second RKHS using a conditionally centered kernel. This yields an RKHS-valued residual moment whose squared norm forms the test statistic. We establish a weighted chi-square null limit and consistency against fixed alternatives for the population-centered statistic. An empirically centered version is shown to retain the null limit and enables spectral calibration of critical values and $p$-values without resampling. Simulation studies demonstrate accurate size control and power against nonlinear alternatives, and real-data applications illustrate the usefulness of the method for detecting nonlinear predictive relationships in time series.

    https://arxiv.org/abs/2609.22007


    Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation

    oai:arXiv.org:2609.19122v2

    arXiv:2609.19122v2 Announce Type: cross Abstract: Visual signals require compact yet sufficient representations for robust downstream prediction. Convolutional sparse coding (CSC) provides an explicit mechanism for suppressing redundant components while preserving signal content, but its sparsity coefficient is typically fixed and manually selected. We propose a training-adaptive convolutional sparse coding framework for robust visual signal representation. Specifically, we unfold the CSC optimization with the Fast Iterative Shrinkage-Thresholding Algorithm (FISTA) and treat the sparsity coefficient as a differentiable variable jointly learned with the network parameters. From the information bottleneck perspective, this coefficient controls the trade-off between information retention and compression: the sparsity term promotes compact representations, while the reconstruction term together with task loss preserves task-relevant signal content. We further introduce a label-free post-training strategy that adjusts the compression strength for corrupted inputs with the main network parameters fixed. Experiments on CIFAR and ImageNet demonstrate competitive clean-data recognition and greatly improved robustness under different input perturbations.

    https://arxiv.org/abs/2609.19122


    Sparse Priors for Efficient Distribution Learning

    oai:arXiv.org:2609.20883v1

    arXiv:2609.20883v1 Announce Type: cross Abstract: Despite the widespread use and success of generative AI techniques today, theoretical guarantees on learning a distribution supported in $d$ dimensions from $n$ samples degrade as $O(n^{-1/\Theta(d)})$, though shown to be minimax optimal. We hypothesize that present bounds are too pessimistic because smoothness assumptions are not enough to capture the structure of distributions that often appear in real applications. Consequently, we introduce the class of sparse priors and define the "Sparse Dimension" as a measure of sparsity of a prior over the space of all distributions. We show that distribution learning under a $k$-sparse prior achieves a Bayesian risk lower bound of $\Omega(\sqrt{k/n})$ under common distance metrics, and show a matching (up to logarithmic terms asymptotically in $n,k$) upper bound for the TV distance under mild additional assumptions. We show the statistical equivalence of distribution learning and learning to sample in the Bayesian setting so that our results apply to learning to sample as well. While $k$ can still depend on the dimension $d$, or a notion of intrinsic dimension, our results show that learning under an appropriate prior overcomes the curse of dimensionality with respect to the dependence on $n$.

    https://arxiv.org/abs/2609.20883


    Continuous Delayed-Memory Stochastic Gradient Descent and Continuous-Time Reinforcement Learning from History of Astrophysical Time Series Studies

    oai:arXiv.org:2609.20906v1

    arXiv:2609.20906v1 Announce Type: cross Abstract: Quasars are luminous objects in the universe that exhibit stochastic brightness variations encoding information about the supermassive black holes powering them, and modeling these variations from ground-based survey data time series, known as light curves, is a statistical challenge. This paper reviews how stochastic differential equations (SDEs) have been adapted with neural network parameterizations to overcome this challenge in history. We create the Continuous-Delayed-Memory Stochastic Gradient Descent which depend on the past state of the discrete iteration process. We performed the simulation on some 2-dimensional landscape and observed some wider-exploration and more precise convergent behavior compared to Vanilla SGD by adjusting hyperparameters. Besides, we proposed a reinforcement learning structure with continuous time policy gradients for exploratory policies without solving HJB PDE, and we show that its optimality conditions recover the Gibbs policy of previous works.

    https://arxiv.org/abs/2609.20906


    Total Variation Distance between Product Distributions: an Analytic Proxy

    oai:arXiv.org:2609.21049v1

    arXiv:2609.21049v1 Announce Type: cross Abstract: We characterize, up to universal constants, the total variation distance between finite products of arbitrary probability measures by a simple formula. A theorem of Latala reduces this expression to one scalar equation. As an application, we characterize the sample complexity of equal-prior binary hypothesis testing uniformly in the weak-detection regime, where squared Hellinger distance alone does not determine the answer, and also characterize the TV distance between multinomial distributions.

    https://arxiv.org/abs/2609.21049


    Layerwise Decoupling for Stable Structured Sparsification of Fully Connected Layers

    oai:arXiv.org:2609.21126v1

    arXiv:2609.21126v1 Announce Type: cross Abstract: We propose a decoupled, layerwise method for structurally sparsifying the fully connected layers of pretrained neural networks. Rather than penalizing all layers jointly, our approach extracts shallow two-layer subnetworks, normalizes the inner weights, and applies a structured group penalty to the outer weight matrix of each block, processing layers sequentially to prune neurons and reduce the width of each layer. We prove that the constrained decoupled objective is equivalent at optimality to a specific joint penalty on the inner and outer weights, for any positively homogeneous activation, and thus admits a clean projected and proximal formulation. Our central finding is that this decoupled reformulation is more robust than coupled methods. In numerical experiments it provides a wider usable range of the regularization strength and a lower rate of catastrophic over-pruning than the tested joint baseline while maintaining comparable accuracy. We establish these properties in controlled classification and sparse-recovery studies, and examine their scope in a high-dimensional PINN stress test and in the feed-forward layers of OPT-1.3B.

    https://arxiv.org/abs/2609.21126


    Reliability-Centered Evaluation of Sparse Longitudinal CT Lesion-Size Forecasting with Conformal Interval Calibration and Gompertz-Inspired Regularization

    oai:arXiv.org:2609.21197v1

    arXiv:2609.21197v1 Announce Type: cross Abstract: Sparse longitudinal CT follow-up limits lesion-size forecasting when only a few prior observations are available. We constructed a five-visit DLT-derived same-lesion trajectory benchmark from DeepLesion and Deep Lesion Tracker (DLT), yielding 205 trajectories from 129 patients. We compared an exploratory conventional sparse-to-final analysis with a primary fixed visit-index horizon design predicting the common log change from T3 to T4 while progressively adding earlier observations, evaluating predictive accuracy, uncertainty reliability, post-hoc conformal interval calibration, subgroup performance, and Gompertz-inspired trajectory regularization. The evaluated methods showed partially overlapping point-prediction accuracy but distinct uncertainty behavior. Mean held-out RMSE across ten training seeds was 0.4726, 0.4305, 0.4499, and 0.4513 for m = 1, 2, 3, 4, indicating the lowest mean RMSE at m = 2; additional history did not improve RMSE. At m = 4, raw Cohort-Level Feature GP coverage was near the 95% nominal level, whereas MC Dropout, Deep Ensemble, and residual-scale intervals were conservative. Patient-level conformal calibration generally produced near-nominal or conservative coverage at the cost of wider intervals. Patient-grouped development cross-validation selected lambda* = 0 for the Gompertz-inspired term. A global population reference frequently opposed lesion-level change directions, and prediction difficulty varied across anatomical subgroups. Overall, additional historical observations provided limited predictive benefit once the prediction horizon was controlled, while predictive accuracy, uncertainty reliability, and trajectory consistency did not necessarily improve together, and should be evaluated jointly in sparse longitudinal imaging.

    https://arxiv.org/abs/2609.21197


    Summary Indices in Treatment Effect Estimation

    oai:arXiv.org:2609.21393v1

    arXiv:2609.21393v1 Announce Type: cross Abstract: This paper studies the practice of combining multiple outcomes into a summary index to estimate a causal effect. For common estimators and index constructions, the estimate equals a weighted sum of the estimated effects on the components, with weights that are implicit and rarely reported. The paper derives the weights and shows that, for inverse-covariance-weighted indices, they can be negative and unrestricted in magnitude, so the index effect can have the opposite sign to every component effect. The paper proposes two procedures for valid inference on the index effect: a variance estimator that accounts for the data-dependent weights, and a shifted t-test that requires no such correction. Conventional t-tests of the null of no effect remain valid. Contrary to common claims, summary indices do not generally improve power. Three published studies illustrate the results.

    https://arxiv.org/abs/2609.21393


    Multi-Domain Clustering via Measure Quantization

    oai:arXiv.org:2609.21664v1

    arXiv:2609.21664v1 Announce Type: cross Abstract: Clustering is a fundamental task in data analysis, typically addressed through centroid-based methods such as K-means. In this work, we present a general framework for multi-domain clustering via measure quantization: given samples from multiple domains, we learn a shared set of cluster prototypes by minimizing a probability metric, such as the Sinkhorn divergence or the Maximum Mean Discrepancy, between each domain's probability measure and the measure of prototypes. Data points are then assigned to clusters either via nearest centroid, or via optimal transport, a collaborative strategy that couples all samples within a domain. A mini-batch optimization strategy makes both fitting and assignment scalable, reducing memory and computational cost while preserving clustering performance. Experimental results on 5 multi-domain benchmarks spanning image, audio and sensor data show that our Sinkhorn-based method consistently outperforms classical and multi-domain clustering baselines, and that this advantage persists when scaling to hundreds of thousands of samples.

    https://arxiv.org/abs/2609.21664


    Match forecasts in UEFA club competitions: Elo ratings versus Transfermarkt valuations

    oai:arXiv.org:2609.21674v1

    arXiv:2609.21674v1 Announce Type: cross Abstract: The pre-season strengths of European football clubs are usually measured by two proxies in the literature. Football Club Elo Ratings provide strictly performance-based Elo ratings from the early days of the European Cups, while Transfermarkt valuations are crowd-based estimates of squad market values. This paper compares them by evaluating their ability to forecast the results of matches played in the UEFA Champions League and the UEFA Europa League between the seasons 2020/21 and 2024/25. The two indicators yield almost identical out-of-sample accuracy when used separately. Combining the two measures leads to a modest improvement, but the best aggregation procedure is sensitive to the forecast target. Our results suggest that seeding based on Elo ratings would be (closely) optimal.

    https://arxiv.org/abs/2609.21674


    A Framework to Quantify the Probability of Future Cyber Loss Events

    oai:arXiv.org:2609.21717v1

    arXiv:2609.21717v1 Announce Type: cross Abstract: Cybersecurity risk quantification remains challenging due to limited operational data and difficulties in quantifying Loss Event Frequency (LEF). This paper introduces the Loss Event Frequency Security Analyser (LEFSA), a probabilistic framework that reformulates LEF estimation as machine-level Cyber Loss Event (CLE) prediction combined with hierarchical infrastructure-level aggregation. LEFSA estimates calibrated machine-level CLE probabilities from operational cybersecurity telemetry and aggregates them across infrastructure layers while accounting for machine-level dependencies. This provides a foundation for scalable, explainable, and operationally applicable cyber risk estimation at the level of machines, services, business processes, and the entire organization. The framework was evaluated using proprietary Managed Detection & Response telemetry from 23 organizations using Microsoft Defender for Endpoint. XGBoost achieved the strongest predictive performance, with a mean area under the receiver operating characteristic curve of 0.90 and consistently low calibration error across evaluation periods. The results demonstrate that operational cybersecurity telemetry contains substantial predictive information for future CLE occurrence, supporting probabilistic machine-level modeling and hierarchical aggregation as a promising foundation for quantitative, data-driven cyber risk management.

    https://arxiv.org/abs/2609.21717


    Single-Loop Stochastic Projected Damped Extragradient Methods for Stochastic Nonconvex--(Strongly) Concave Minimax Optimization

    oai:arXiv.org:2609.21747v1

    arXiv:2609.21747v1 Announce Type: cross Abstract: We develop single-loop stochastic projected damped extragradient methods for stochastic nonconvex--(strongly) concave minimax optimization, with complexity guarantees for both game stationarity (GS) and optimization stationarity (OS). Our approach combines a stochastic projected damped extragradient (SPDE) method with a recursive variance-reduced variant, VR-SPDE, both of which retain a single-loop structure. Under an unbiased stochastic gradient oracle with uniformly bounded variance, SPDE finds an $\varepsilon$-game-stationary point with stochastic first-order oracle (SFO) complexities of $O(\kappa\varepsilon^{-4})$ and $O(\varepsilon^{-5})$ in the nonconvex--strongly concave and nonconvex--concave settings, respectively, where $\kappa=L/\mu$. Under an additional mean-square Lipschitz condition on the stochastic gradients, VR-SPDE improves these GS complexities to $O(\kappa^{3/2}\varepsilon^{-3})$ and $O(\varepsilon^{-9/2})$, respectively. For an $\varepsilon$-optimization-stationary point, SPDE achieves SFO complexities of $O(\kappa\varepsilon^{-4})$ and $O(\varepsilon^{-6})$, while VR-SPDE achieves $O(\kappa^{3/2}\varepsilon^{-3})$ and $O(\varepsilon^{-6})$, in the two settings, respectively. These OS guarantees match the best-known bounds achieved by multi-loop methods while preserving a single-loop implementation. To the best of our knowledge, our results provide the best-known SFO complexity guarantees among single-loop stochastic first-order methods for the respective stationarity criteria and problem classes.

    https://arxiv.org/abs/2609.21747


    EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise

    oai:arXiv.org:2609.21841v1

    arXiv:2609.21841v1 Announce Type: cross Abstract: Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically valuable tasks, yet most enterprise GenAI initiatives fail to show a measurable business effect and a large fraction of agentic projects are expected to be cancelled. We argue that this is substantially a measurement problem: public benchmarks answer "what can the model do?", whereas a deployment decision requires "is this workflow fit, reliable, safe and worth scaling - here, on our data, under our controls?". We present EnterpriseVal, a use-case-level evaluation system that closes this gap. It comprises (i) a formal specification of the use case and of the frozen socio-technical configuration under test, model, prompts, retrieval, tools, guardrails and human oversight, with an autonomy level and consequence tier that jointly set the required evaluation intensity; (ii) a metric catalogue spanning fidelity, utility, efficiency, reliability, assurance and oversight; (iii) a grading protocol that scales blinded expert judgement with calibrated LLM-as-judge scoring through prediction-powered inference; (iv) a two-tier threshold gate, stated as an executable algorithm, that maps metric vectors with confidence bounds to REJECT/CONDITIONAL/SCALE decisions; and (v) a value-and-risk model in which the reviewer catch rate is a measured parameter. We report a pilot across three workflows in a global bank. In credit-memo drafting, human-graded citation precision reached 88% and hallucination rate 1.6% for the best model against gates of 70% and 5%; in procedure transformation, analyst refinement effort fell from an estimated 27.4 to 2.9 hours per document. We separate established results, documented pilot evidence, the proposed system and open hypotheses, and specify the experiments required for full validation

    https://arxiv.org/abs/2609.21841


    Geometric Mean Pooling for Equal-Weight Multiplicative Coarse-Graining

    oai:arXiv.org:2609.21876v1

    arXiv:2609.21876v1 Announce Type: cross Abstract: As an alternative to the additive and extremal biases of average and max pooling, we introduce Geometric Mean Pooling (GMP), a signed pooling operator that combines the product of feature signs with the geometric mean of feature magnitudes. Motivated by local-to-global composition in quantum many-body physics, GMP retains both joint sign information and a characteristic multiplicative scale without introducing learnable pooling parameters. We show that non-overlapping hierarchical GMP preserves the corresponding global multiplicative statistic and evaluate it on synthetic sequence tasks, iterative coarse-graining, image classification, and molecular lipophilicity regression. On the synthetic tasks, GMP recovers product-based signals more accurately than average and max pooling and maintains predictive performance under the tested levels of multiplicative input noise. On image and molecular data, however, its effectiveness depends on the representation, target parameterization, and placement of local and global pooling. These results position GMP as a complementary, regime-dependent inductive bias for tasks in which equal-weight multiplicative composition is plausible, rather than as a universal replacement for standard pooling operators.

    https://arxiv.org/abs/2609.21876


    Interval-Constrained Brownian Paths: Exact Interpolation and Extrapolation

    oai:arXiv.org:2609.21943v1

    arXiv:2609.21943v1 Announce Type: cross Abstract: We study Brownian motion and Brownian bridge processes conditioned to remain in a fixed interval $[0,a]$, focusing on the conditional distribution at a single time. For a Brownian motion in $[0,a]$, conditioned on survival up to time $t$ we recall (and present in a self-contained form) the conditional density of its position at $t$ (exact extrapolation). For a Brownian bridge conditioned to remain in $[0,a]$ on $[0,T]$ we express the probability density at an interior time as a normalized product of killed transition densities (exact interpolation). These densities admit dual complementary series representations, which are linked via the Jacobi theta identity. Our main contribution is a unified exact sampling suite for both extrapolation and interpolation that (i) includes boundary endpoints and (ii) automatically switches between density representations to keep acceptance rates efficient across regimes. To that end, we derive simple proposal families which cover both small- and large-time regimes, with an automatic rule selecting the tighter envelope. Our procedures extend to exact simulation of discrete skeleton paths at multiple times via the Markov property.

    https://arxiv.org/abs/2609.21943


    Moral Entropy: Auditing Bias and Uncertainty in Moral Judgment

    oai:arXiv.org:2609.21992v1

    arXiv:2609.21992v1 Announce Type: cross Abstract: Most work in computational ethics treats annotator disagreement on moral content as noise to be voted away, collapsed into majority vote or the more permissive any-annotator rule the moment a single annotator flags an item. We argue this uncertainty should instead be modeled and learned from. We introduce Moral Entropy, a Bayesian framework that keeps a full posterior over the true label and decomposes its entropy into aleatoric uncertainty (irreducible disagreement about the moral content) and epistemic uncertainty (from insufficient or noisy annotation) -- and lets any heuristic consensus rule be audited against a calibrated ground truth via entropy methods such as cross-entropy/KL, Brier score, and expected calibration error. Across three corpora and fifteen discourse domains, auditing the standard aggregation rules against this posterior reveals bias that no current pipeline reports: the any-annotator rule disagrees with the calibrated posterior on roughly 30% of items -- pooled, almost entirely false positives, though the errors invert at the foundation level (19.9%/38.9% mean FPR/FNR on MFTC) -- while the stricter majority and two-vote rules miss 63-83% of true positives.

    https://arxiv.org/abs/2609.21992


    Auditing bipartite motif interpretations: a worked example with conservation checks and open-path decomposition

    oai:arXiv.org:2609.22014v1

    arXiv:2609.22014v1 Announce Type: cross Abstract: Motif profiles of bipartite agent-object networks, such as tourist-site visits and customer-item transactions, are read as evidence about structural roles and about differences between networks, often without asking what the two degree sequences already fix. In a simple bipartite graph the induced k-fan count on one node type is a sum of degree combinations, so it has zero variance under a null that preserves both degree sequences. We apply this known result to a reconstructed tourism rating network of 17 tourists, 80 sites and 637 edges, the sole inferential worked example, and, as a provenance-limited illustration, to published motif-instance aggregates over 36 monthly luxury customer-item networks. The four fan classes are exact functions of the degree sequences: in the tourism network the raw fan counts and the size-3 two-fan ratio (84.7% fan-out) restate those sequences. The published luxury counts require at least 89,502 customer-item edges against 26,451 reported transactions, so their 99.8% fan-in is reported as a descriptive value only. Against a hard bipartite configuration null, the four-cycle count is degree-consistent (z about +1.0) and the open path is deficient (z about -5.8) by 2,243 instances, 2.4% of the null mean; the deficit survives every leave-one-tourist-out re-run (z -4.5 to -8.1). An exact identity splits it at the point estimate into 64.4% mixing and 35.6% four-cycle, but that split is not an attribution: the observed mixing term lies below all 500 null samples, the two components are almost collinear under the null (r = 0.968), and the mixing share ranges from 36.6% to 116.5% under leave-one-tourist-out deletion. The deficit is extreme relative to the sampled null, while its class-level interpretation is undetermined and unstable. We give a four-step pre-interpretation check and a reference implementation.

    https://arxiv.org/abs/2609.22014


    Optimized variance estimation under interference and complex experimental designs

    oai:arXiv.org:2112.01709v3

    arXiv:2112.01709v3 Announce Type: replace Abstract: Unbiased and consistent variance estimators generally do not exist for design-based treatment effect estimators because experimenters never observe more than one potential outcome for any unit. The problem is exacerbated by interference and complex experimental designs. Experimenters must accept conservative variance estimators in these settings, but they can strive to minimize the conservativeness. In this paper, we show that the task of constructing a minimally conservative variance estimator can be interpreted as an optimization problem that aims to find the lowest estimable upper bound of the true variance given the experimenter's risk preferences and knowledge of the potential outcomes. We characterize the set of admissible bounds in the class of quadratic forms, and we demonstrate that the optimization problem is a convex program for many natural objectives. The resulting variance estimators are guaranteed to be conservative regardless of whether the background knowledge used to construct the bound is correct, but the estimators are less conservative if the provided information is reasonably accurate. Numerical results show that the resulting variance estimators can be considerably less conservative than existing estimators, allowing experimenters to draw more informative inferences about treatment effects.

    https://arxiv.org/abs/2112.01709


    Targeted Function Balancing: Flexible balancing weights that leverage strategic imbalance to estimate causal effects

    oai:arXiv.org:2203.12179v4

    arXiv:2203.12179v4 Announce Type: replace Abstract: When estimating the average treatment effect of a binary intervention with weights, investigators are often burdened with how best to leave imbalance in the distributions of observed covariates. This requires that investigators proceed cautiously, as imbalance in key (nonlinear transformations of) covariates potentially introduces bias into the ultimate estimated effect. This paper introduces Targeted Function Balancing (TFB), a flexible covariate balancing weights approach that simplifies this problem for the researcher. TFB first regresses an outcome on covariates, and then selects weights that balance functions (of the covariates) that are probabilistically near the resulting regression function. This yields balance in the regression function's predicted values and the covariates, with the regression function's estimated variance determining how much balance in the covariates is sufficient. Notably, TFB demonstrates that intentionally leaving imbalance in some covariates can increase efficiency without introducing bias, challenging traditions that warn against imbalance in any variable. Though, TFB still does prioritize balance in the covariates that are most influential on the outcome. Additionally, TFB is entirely defined by a regression function and its estimated variance, turning the problem of how best to balance the covariates into how best to model the outcome. Kernel regularized least squares, the LASSO, and Bayesian Additive Regression Trees are considered as regression estimators. We apply TFB to estimate the effects of taking mathematics courses in twelfth grade in the Los Angeles Unified School District. The R Package tfb implements TFB.

    https://arxiv.org/abs/2203.12179


    Bias-Correction for Privacy-Protected Spatial Autoregressive Models with Application to Restaurant Network Analysis

    oai:arXiv.org:2403.16773v3

    arXiv:2403.16773v3 Announce Type: replace Abstract: Spatial autoregressive (SAR) models and their extensions are important tools for studying network effects. However, with an increasing emphasis on data privacy, data providers often implement protection measures that render standard SAR models inapplicable. In this study, we introduce a privacy-protected SAR model that incorporates noise into both the response and covariates to meet privacy requirements. With noise present in both components, the traditional quasi-maximum likelihood estimator becomes difficult to compute because the likelihood function cannot be directly formulated. To bypass this hurdle, we begin with a pseudo-likelihood approach, initially omitting the noise in the covariates. A Newton-Raphson algorithm is then applied to compute the estimator; however, the estimator is biased. To address this, we propose a bias-corrected Newton-Raphson-type algorithm that simultaneously accounts for noise in both the response and covariates. We further show, under appropriate regularity conditions, that the resulting estimator is consistent and asymptotically normal. To further enhance computational efficiency, we also develop a bias-corrected least squares estimator. Several extensions are discussed, and the finite-sample performance of the proposed methods is evaluated through extensive simulations. We apply the proposed methodology to restaurant transaction data from a third-party payment platform. Our method identifies a statistically significant competitive network effect among restaurants and further reveals meaningful restaurant-customer interaction patterns.

    https://arxiv.org/abs/2403.16773


    Comparing multilevel and fixed effect approaches in the generalized linear model setting

    oai:arXiv.org:2411.01723v3

    arXiv:2411.01723v3 Announce Type: replace Abstract: We extend prior work comparing linear multilevel models (MLM) and fixed effect (FE) models to the generalized linear model (GLM) setting, where the coefficient on a treatment variable is of primary interest. This leads to three insights. (i) First, as in the linear setting, MLM can be thought of as a regularized form of FE (RegFE). This explains why group-level confounding can greatly bias MLM's treatment coefficient estimates. However, unlike the linear setting, there is not an exact equivalence between MLM, fit through (approximate) maximum likelihood estimation, and RegFE in GLMs. (ii) Second, we study a generalization of "bias-corrected MLM" (bcMLM) to the GLM setting, and a corresponding "bias-corrected RegFE" (bcRegFE). None of FE, bcMLM, or bcRegFE entirely solve MLM's bias problem in GLMs, but bcMLM and bcRegFE tend to show less bias than does FE. (iii) Third, as in the linear setting, MLM's default standard errors can misspecify the true intragroup dependence structure in the GLM setting, which can yield downwardly biased standard errors. A cluster bootstrap is a more agnostic alternative. We also consider a cluster-robust standard error for (bc)RegFE. Ultimately, for non-linear GLMs, we recommend bcMLM for estimating the treatment coefficient, and a cluster bootstrap for standard errors and confidence intervals. If a bootstrap is not computationally feasible, then we recommend bcRegFE with cluster-robust standard errors, or FE with cluster-robust standard errors when group sizes are larger. We apply our recommendations to estimate the effects of taking a mathematics course in twelfth grade for students in the Los Angeles Unified School District. Our mlmfe package easily implements our recommendations in the R programming language.

    https://arxiv.org/abs/2411.01723


    Functional independent component analysis by choice of norm: a framework for near-perfect classification

    oai:arXiv.org:2412.17971v3

    arXiv:2412.17971v3 Announce Type: replace Abstract: We develop a theory for functional independent component analysis in an infinite-dimensional framework using Sobolev spaces that accommodate smoother functions. The notion of penalized kurtosis is introduced motivated by Silverman's method for smoothing principal components. This approach allows for a classical definition of independent components obtained via projection onto the eigenfunctions of a smoothed kurtosis operator mapping a whitened functional random variable. We discuss the theoretical properties of this operator in relation to a generalized Fisher discriminant function and the relationship it entails with the Feldman-H\'ajek dichotomy for Gaussian measures, both of which are critical to the principles of functional classification. The proposed estimators are a particularly competitive alternative in binary classification of functional data and can eventually achieve the so-called near-perfect classification, which is a genuine phenomenon of high-dimensional data. Our methods are illustrated through simulations, various real datasets, and used to model electroencephalographic biomarkers for the diagnosis of depressive disorder.

    https://arxiv.org/abs/2412.17971


    Online activity prediction via generalized Indian buffet process models

    oai:arXiv.org:2505.19643v3

    arXiv:2505.19643v3 Announce Type: replace Abstract: Online A/B tests are the standard tool for data-driven decision-making at scale. Among the design choices with the largest impact on statistical power is the triggering mechanism: how many users to expose and for how long. This often requires forecasting user engagement, i.e., whether enough users will trigger, and when a target participation level will be reached, from limited pilot data. We introduce a Bayesian nonparametric model for predicting both new-user counts and total triggers, accommodating the heavy-tailed engagement patterns typical of web experiments. All predictive quantities can be computed without intensive numerical procedures such as Markov chain Monte Carlo (MCMC) or variational inference. We evaluate on three public datasets (over 450 public benchmark evaluations) and a proprietary benchmark drawn from 759 production A/B tests comprising 1,774 arms. Across the benchmark analyses, our models are competitive and frequently improve accuracy in forecasting new users, total triggers, and time to reach a target sample size compared with state-of-the-art competitors, especially when only a few pilot days are observed.

    https://arxiv.org/abs/2505.19643


    Doubly robust integration of nonprobability and probability survey data

    oai:arXiv.org:2508.05859v3

    arXiv:2508.05859v3 Announce Type: replace Abstract: Doubly robust (DR) estimators of the population mean of an outcome have been proposed for integrating outcome and covariate data from a nonprobability sample with covariate data from a probability survey. These estimators combine inverse probability weighting with mass imputation. We review these DR estimators and extend them to allow estimation of a domain (subpopulation) mean, possibly using data from individuals outside the domain to improve estimation when the domain is small. For situations where an outcome is observed in both samples, we describe estimators that efficiently combine one of these DR estimators with a Horvitz-Thompson or Hajek estimator that uses only the probability survey data. We investigate the asymptotic relative efficiency of these combined estimators compared to their two component estimators, and carry out a simulation study to assess relative efficiency in finite samples. The relative efficiency depends on the ratio of the variances of the two component estimators and on how predictive the covariates are of the outcome. We illustrate the use of all the estimators by analysing data from the Natsal-3 probability survey and four contemporaneous online nonprobability surveys. Our extensions will enable greater uptake and practical use of these combination methods in the survey statistics field.

    https://arxiv.org/abs/2508.05859


    Robust Mixture Models for Algorithmic Fairness Under Latent Heterogeneity

    oai:arXiv.org:2509.17411v2

    arXiv:2509.17411v2 Announce Type: replace Abstract: Machine learning models optimized for average performance can perform poorly on vulnerable subpopulations. Existing approaches often rely on groups specified in advance, yet fairness-relevant subgroup structure may be latent, intersectional, and driven by complex interactions among continuous and discrete attributes. We introduce \textbf{ROME} (\textbf{\underline{RO}}bust \textbf{\underline{M}}ixture \textbf{\underline{E}}nsemble), a framework that learns latent group structure while optimizing worst-group predictive performance. ROME connects latent-variable modeling with distributionally robust optimization (DRO) through two complementary approaches: an Expectation-Maximization formulation with robust aggregation for linear models and a neural Mixture-of-Experts formulation for nonlinear settings. Across simulations and three real-world regression datasets, ROME improves worst-group performance while maintaining competitive overall accuracy, including in comparisons with established group-aware and group-label-free robust learning methods. ROME provides a flexible approach to robust prediction when fairness-relevant attributes are available for subgroup discovery but their direct use in group-specific outcome models is restricted.

    https://arxiv.org/abs/2509.17411


    Scalable Bayesian Spatiotemporal Tensor Modeling for Censored and Missing Areal Data

    oai:arXiv.org:2511.17725v4

    arXiv:2511.17725v4 Announce Type: replace Abstract: We propose a new Bayesian approach for spatiotemporal areal data with censored and missing observations, where the response is represented as a tensor indexed by spatial units and time points. The method introduces a flexible random effect that combines the spatial dependence structures of the Simultaneous Autoregressive (SAR) and Directed Acyclic Graph Autoregressive (DAGAR) models with a temporal autoregressive component. This formulation brings both spatial models into a unified spatiotemporal framework by expressing them as Gaussian Markov random fields in innovation form, providing an interpretable representation of spatial, temporal, and joint spatiotemporal dependence while preserving the multiway structure of the data. The innovation-based structure enables scalable Bayesian implementation in \texttt{Stan} for datasets of moderate size. Simulation studies show that the proposed model outperforms common ad hoc imputation strategies for censored and missing data. We further illustrate its practical value through an application to carbon monoxide (CO) concentrations, where the proposed framework captures complex spatiotemporal dependence in the presence of censoring and missingness.

    https://arxiv.org/abs/2511.17725


    Randomized Principal Component Ensembles for High-Dimensional Calibration

    oai:arXiv.org:2512.09505v2

    arXiv:2512.09505v2 Announce Type: replace Abstract: Calibration is a widely used method in survey sampling to adjust weights so that estimated totals of some chosen calibration variables match known population totals or totals obtained from other sources. When a large number of auxiliary variables are included as calibration variables, the variance of the total estimator can increase, and the calibration weights can become highly dispersed. To address these issues, we propose an ensemble method based on a principal component decomposition of the auxiliary variables. We repeatedly select sets of principal components without replacement and with unequal selection probabilities. Calibration is performed on each selected set of principal components, and the resulting weights are aggregated to obtain a final set of weights. With our proposed method, it is possible to calibrate exactly for some of the main auxiliary variables, while relaxing the calibration constraints for the remaining variables. It yields a total estimator whose variance does not explode when new auxiliary variables are added while producing weights with low dispersion. Finally, our proposed method allows us to obtain a single weighting system that can be applied to several variables of interest of a survey. An estimator of the variance of the total estimator is also proposed. We evaluate the proposed total estimator and its variance estimator using a simulation study on real survey data from the Swiss Survey on Income and Living Conditions and on synthetic data. The results show that the proposed solution significantly reduces the weight variability and the variance of the total estimator compared with competing total estimators for some variables of interest.

    https://arxiv.org/abs/2512.09505


    Boltzmann generators for amorphous particle systems

    oai:arXiv.org:2512.16607v3

    arXiv:2512.16607v3 Announce Type: replace Abstract: Sampling configurations in thermodynamic equilibrium is a long-standing challenge in statistical physics. Boltzmann generators address this problem by employing generative models to propose independent configurations, which are then reweighted via importance sampling using exact likelihood evaluations. Recent Boltzmann Generators based on continuous normalizing flows and flow matching have achieved significant success for particle systems and biomolecules. However, these approaches have not been extended to amorphous materials (glasses), for which equilibrium sampling is notoriously slow. Because of their disordered structure, the invariances and geometrical constraints of amorphous materials differ from those of crystals and biomolecules, preventing the direct use of existing generative models. Here, we develop Boltzmann Generators tailored to amorphous materials by building the required equivariances directly into Riemannian stochastic interpolants. Our framework incorporates periodic boundary conditions and particle symmetries using equivariant graph neural networks. Numerical experiments demonstrate that enforcing physical symmetries significantly improves the accuracy of Boltzmann Generators, but also reveal an intrinsic limitation of the continuous-flow formulation: accumulated numerical errors during likelihood integration break time-reversibility, compromising exact thermodynamic reweighting. These results reveal a fundamental challenge for continuous-flow generative models in statistical mechanics and call for alternative approaches that preserve exact thermodynamic consistency.

    https://arxiv.org/abs/2512.16607


    Delayed Acceptance Slice Sampling

    oai:arXiv.org:2512.17868v2

    arXiv:2512.17868v2 Announce Type: replace Abstract: Slice sampling is a well-established Markov chain Monte Carlo method for approximate sampling of target distributions which are only known up to a normalizing constant. The method is based on choosing a new state on a slice, i.e., a superlevel set of the given unnormalized target density (with respect to a reference measure). However, slice sampling algorithms usually require per step multiple evaluations of the target density, and thus can become computationally expensive. This is particularly the case for Bayesian inference with costly likelihoods. In this paper, we exploit deterministic approximations of the target density, which are relatively cheap to evaluate, and propose delayed acceptance versions of several common (hybrid) slice samplers. We show ergodicity of the resulting slice sampling methods, discuss the superiority of delayed acceptance (ideal) slice sampling over delayed acceptance Metropolis-Hastings algorithms, and illustrate the benefits of our novel approach in terms of improved computational efficiency in numerical experiments.

    https://arxiv.org/abs/2512.17868


    Near-Universal Multiplicative Updates for Nonnegative Einsum Factorization

    oai:arXiv.org:2602.02759v3

    arXiv:2602.02759v3 Announce Type: replace Abstract: Despite the ubiquity of multiway data across scientific domains, there are few performant and user-friendly methods that fit non-standard nonnegative tensor factorization models tailored to the data at-hand. Researchers may use gradient-based automatic differentiation, which often struggles under nonnegative constraints, choose between a limited set of methods with mature implementations, or implement their own model from scratch. As an alternative, we introduce NNEinFact, an einsum-based multiplicative update algorithm that fits any nonnegative tensor factorization expressible as a tensor contraction by minimizing one of many user-specified loss functions, including the $(\alpha,\beta)$-divergence. To use NNEinFact, the researcher specifies their model with a string. NNEinFact converges to a stationary point of the loss, supports missing data, and fits to tensors with hundreds of millions of entries in seconds. Empirically, NNEinFact fits custom models which outperform standard ones in prediction tasks on real-world tensor data by over $37\%$ and attains less than half the test loss of gradient-based methods while converging up to 90 times faster. Software is publicly available at https://github.com/jhood3/einfact.

    https://arxiv.org/abs/2602.02759


    A variational framework for modal estimation

    oai:arXiv.org:2602.17956v2

    arXiv:2602.17956v2 Announce Type: replace Abstract: Multivariate mode estimation arises in many statistical problems such as inverse problems, multimodal sampling, and density-based clustering, but becomes challenging in moderate to high dimensions, especially when the underlying density is not directly evaluable. We introduce GERVE (Gibbs-measure Entropy-Regularized Variational Estimation), a sample-based method for estimating multivariate modes by approximating Gibbs distributions directly from samples, without estimating or evaluating the density. GERVE uses Gaussian-mixture variational annealing and natural-gradient optimization, producing a mixture concentrated in high-density regions whose component responsibilities also provide a clustering of the observations. We prove theoretical guarantees in two regimes: as the Gibbs temperature goes to zero, the optimal variational mixture concentrates around the global modes of the population density; at fixed positive temperature, we prove existence, consistency, and asymptotic normality of empirical maximizers and propose a bootstrap procedure for uncertainty quantification. Simulations and a real-data experiment show that GERVE accurately recovers modes and produces meaningful clusters.

    https://arxiv.org/abs/2602.17956


    Objective Model Prior Probabilities in Variable Selection

    oai:arXiv.org:2603.19728v2

    arXiv:2603.19728v2 Announce Type: replace Abstract: For many years it was routine to use equal model prior probabilities in Bayesian model uncertainty analysis. At least twenty years ago it became clear that this was problematic, leading to support of much too large models in the increasingly huge model spaces being considered in genomics and other fields. A popular replacement was to adopt a suggestion of Harold Jeffreys for the variable selection problem in which a total of $k$ possible variables are being considered for inclusion in the model: give the collection of all models containing $d$ variables ($d = 0, . . . , k$) prior probability $1/(k + 1)$ and then divide this prior probability equally among the models in the collection. Many other choices of model prior probabilities that impose severe parsimony have also been introduced. We begin by reviewing the problems with using equal model prior probabilities and then discuss some serious problems with the Jeffreys choice. Finally, we introduce and study a number of objective alternative choices of model prior probabilities, from both numerical and theoretical perspectives.

    https://arxiv.org/abs/2603.19728


    Why is Regularization Underused? An Empirical Study on Trust and Adoption of Statistical Methods

    oai:arXiv.org:2604.02992v2

    arXiv:2604.02992v2 Announce Type: replace Abstract: Statistical practice does not automatically follow methodological innovation. Regularization methods, widely advocated to reduce overfitting and stabilize inference, are readily available in modern software, but are not consistently used by data analysts. We investigate this implementation gap in a large-scale empirical study of trust in, and acceptance of, regularization techniques, based on $N = 606$ data analysts. Drawing on measurement frameworks from technology acceptance research, we survey practitioners and embed a randomized experiment to test whether written recommendation of regularization methods increases trust or intended use. We find no evidence of such an effect. Instead, adoption intentions are strongly associated with analysts' perceptions of ease of implementation and practical benefit, such as improved bias control or interpretability. Perceived social norms also emerge as a central driver. These results indicate that uptake of statistical methodology depends less on formal recommendations than on usability, perceived utility, and community practice.

    https://arxiv.org/abs/2604.02992


    Amortized Filtering and Smoothing with Conditional Normalizing Flows

    oai:arXiv.org:2604.07169v3

    arXiv:2604.07169v3 Announce Type: replace Abstract: Bayesian filtering and smoothing are central to data assimilation in nonlinear dynamical systems. Recent advances in deep generative models provide flexible approximations of the associated non-Gaussian posterior distributions. However, several existing approaches require the score to be re-estimated or a transport map to be constructed at each assimilation step. We propose an amortized framework for filtering and smoothing that reuses trained conditional models across observation sequences and assimilation times. The framework jointly learns a shared recurrent summary network and two conditional normalizing flows from simulated state and observation trajectories. The recurrent network represents each observation history by a fixed-dimensional summary that conditions the filtering approximation and, together with the next state, the backward transition approximation. An information-theoretic analysis shows that, under the Markov assumption, a summary sufficient for filtering is also sufficient for the backward kernel. Combining the terminal filtering approximation with the backward kernels yields an approximate joint smoothing distribution. Numerical experiments on advection-diffusion, Burgers, and Lorenz systems demonstrate the accuracy of the proposed filtering and smoothing approximations and characterize the evolution of their errors beyond the training horizon.

    https://arxiv.org/abs/2604.07169


    A bivariate cure copula model with zero-inflated gamma frailty for right-censored data

    oai:arXiv.org:2604.23154v2

    arXiv:2604.23154v2 Announce Type: replace Abstract: In biomedical studies, paired survival data arise naturally when two event times are observed within the same subject. Existing statistical models seldom accommodate both cure fractions and complex dependence structures. In this paper, we propose a novel bivariate cure frailty-copula model for paired survival data with a cure fraction. By incorporating a zero-inflated gamma frailty, the proposed framework simultaneously accommodates a cure fraction and continuous unobserved heterogeneity among uncured subjects. Dependence between cure statuses is modeled naturally via an odds-ratio parameter, while dependence between survival times conditional on frailty is captured through a copula. To provide interpretable measures of overall dependence in the presence of cure, we derive population-level rank correlation coefficients, namely tie-adjusted versions of Kendall's tau and Spearman's rho. For suitable choices of baseline marginal survival functions and copulas, the joint survival function admits a closed-form expression, enabling maximum likelihood estimation and likelihood ratio testing. Simulation studies demonstrate satisfactory finite-sample performance and favorable comparisons with existing bivariate cure models. In a real-data analysis, the proposed model is favored by information criteria. An R package, curecopula, implementing the proposed methods is publicly available on GitHub.

    https://arxiv.org/abs/2604.23154


    ISOMORPH: A Supply Chain Digital Twin for Simulation, Dataset Generation, and Forecasting Benchmarks

    oai:arXiv.org:2605.12768v3

    arXiv:2605.12768v3 Announce Type: replace Abstract: Open time-series forecasting (TSF) benchmarks cover retail, energy, weather, and traffic, but supply-chain logistics remains underserved. We introduce ISOMORPH, the first public digital twin of a multi-echelon logistics network with interpretable, user-configurable parameters and modular topology, demand, and control rules. The simulator advances a directed routing graph in discrete time: demand is served from inventory or recorded as backlog and triggers replenishment throughout the network. The state tracks inventory, outstanding orders, in-transit shipments, and a smoothed demand estimate, yielding Markovian dynamics on a tractable state space. The released data reproduces the bullwhip effect at empirically consistent magnitudes, while three conservation laws provide verification tools for simulator extensions. We release datasets at two catalogue scales ($C=50$ and $C=200$), with a 33-rollout scenario library at $C=50$. These datasets exhibit dynamics largely absent from fixed TSF benchmarks, including variance amplification, cascading bottlenecks, regime shifts, and cross-channel coupling through shared macro shocks. Zero-shot evaluation of three foundation models (Chronos, Moirai, TimesFM) against three in-domain-trained baselines (ARIMA, ETS, PatchTST) spans four targets: demand, backlog, fill rate, and edge utilization. Comparison with ETTh1, Electricity, and Weather shows that ISOMORPH introduces forecasting regimes that differ from standard real-world TSF benchmarks, positioning it as a complementary, regenerable logistics-domain benchmark. The same pairing produces forecast confidence bands across scenario configurations, providing forward UQ from parameter uncertainty and demonstrating foundation models as fast surrogates for digital-twin-based UQ. Code (MIT): https://github.com/tuhinsahai/ISOMORPH. Interactive demo: https://huggingface.co/spaces/HyeminGu/ISOMORPH-demo.

    https://arxiv.org/abs/2605.12768


    Nonparametric Inference Based on Expected Order Statistics

    oai:arXiv.org:2605.25897v2

    arXiv:2605.25897v2 Announce Type: replace Abstract: In many statistical procedures, such as smoothing, resampling, or privacy-preserving release, it is convenient to transform the empirical distribution according to the task. We introduce an averaging transformation that assigns equal mass to $m$ estimated expected order statistics, where $m$ acts as a resolution parameter. This yields finite-sample properties and a tractable route to asymptotic analysis. The distance between the transformed empirical distribution and its deterministic counterpart is bounded by the estimation error of the classical empirical distribution, in both Wasserstein and integral metrics. Moreover, every $L$-functional of the transformed distribution corresponds to an $L$-functional of the classical one, with updated weights. We establish almost sure and weak convergence, asymptotic distributions for distance-based functionals, and bootstrap validity. Finally, we obtain a differentially private release with explicit accuracy bounds, and compare it with private histograms in simulations and data.

    https://arxiv.org/abs/2605.25897


    Trace-Class Results for MCMC Algorithms for Student-$t$ Regression Models

    oai:arXiv.org:2606.05672v3

    arXiv:2606.05672v3 Announce Type: replace Abstract: In this paper, we consider MCMC algorithms for Student-$t$ regression models. In three cases, we investigate the efficiency of Markov chains based on the algorithms in terms of whether trace-class results hold or not. First, we consider the case where the parameters follow a matrix-normal-inverse-Wishart distribution and show that the Markov operator associated with a standard data augmentation algorithm is trace-class. Second, we consider the case of an improper prior and univariate outcomes. In this case, the standard Markov operator is not trace-class but the Markov operator associated with a collapsed Gibbs algorithm is trace-class. Third, we consider the case of an improper prior and multivariate outcomes. We obtain a trace-class result for a parameter expanded data augmentation algorithm which is based on a univariate working parameter. Finally, we consider the problem of numerially estimating a convergence rate of the trace-class Markov operator in the second case.

    https://arxiv.org/abs/2606.05672


    A Taxonomy of Distance Metrics for Time-Sensitive Importance Splitting: Timer Bounds, Resampling, and the Global Age

    oai:arXiv.org:2607.17939v2

    arXiv:2607.17939v2 Announce Type: replace Abstract: Importance splitting (ISPLIT) evaluates the probabilities of rare events in non-Markovian models. It requires a heuristic importance function (IFUN) that estimates the distance to the target. While including timer evaluations in the IFUN can substantially improve the effectiveness of ISPLIT, the existing time-sensitive IFUNs evaluate simulation states with respect to single sampled timer values. Thus, reaching highly important states requires simultaneously sampling specific combinations of timer values, yielding many unproductive simulation runs. In this paper, we revisit time-sensitive ISPLIT with the goal of steering simulation runs towards important states. First, we study how timer values can be resampled conditioned on the elapsed time. The importance can be evaluated by considering the set of feasible timer values, decoupling importance estimation from timer samples. Second, we exploit the global age of a simulation to identify and prune the executions that can no longer reach the target within the remaining time budget. Together, these ideas lead to a taxonomy of distance metrics clarifying the role of timer bounds, resampling, and the global age. In particular, for models with unbounded timers, we show that time-sensitive IFUNs collapse to ordinary IFUNs under resampling. Experiments demonstrate that the proposed formulations substantially improve the accuracy of ISPLIT estimators.

    https://arxiv.org/abs/2607.17939


    High-probability guarantees for linear accessibility in feature superposition

    oai:arXiv.org:2609.09556v2

    arXiv:2609.09556v2 Announce Type: replace Abstract: Neural networks can leverage feature superposition to encode more concepts than dimensions, but cross-feature interference constrains the linear accessibility of simultaneously active features. By framing linear accessibility as a compressed sensing problem, we derive high-probability bounds for fixed supports under subgaussian noise, proving the sufficient dimension scales linearly ($d=O_{\varepsilon}(k \log m)$) rather than prior worst-case quadratic limits. We characterize the asymmetry between active and inactive interference and the trade-off between interference and observation-noise budgets. We then validate these bounds across system parameters through Gaussian-tail approximations. We also introduce IHT-SAE, which uses learned iterative refinement to improve feature recovery beyond the limits of linear availability. These results quantify the geometric constraints of the linear representation hypothesis, providing a framework for evaluating sparse autoencoders, compositional generalization, and neural interpretability.

    https://arxiv.org/abs/2609.09556


    Data Attribution via Sketched Metadifferentiation

    oai:arXiv.org:2609.15044v2

    arXiv:2609.15044v2 Announce Type: replace Abstract: Data attribution seeks to quantify how individual training examples shape a model's predictions and underpins problems including data valuation, machine unlearning, and model interpretability. Despite having a long line of work, computationally scalable methods often struggle to predict the effect of removing training data in neural networks due to their non-convex nature. To overcome this challenge, metagradient-based methods such as MAGIC (Ilyas and Engstrom, 2025) differentiate each prediction through the entire training run and compute its exact influence with respect to the training data, but require a separate run for every prediction. To reduce this cost, we cast budgeted attribution as estimating a large influence matrix from a small number of measurements. We show that the measurements most appropriate for recovering this matrix differ from those best suited for attribution itself. We then present two algorithms, MAGE and SPELL, suited for reconstruction and attribution respectively, that run on existing metagradient machinery at no extra cost. Empirical studies demonstrate strong performance over existing baselines across training scales and measurement budgets.

    https://arxiv.org/abs/2609.15044


    Online Bayesian Model Averaging with Joint Uncertainty Quantification for Models and Regression Coefficients in Binary Regression

    oai:arXiv.org:2609.17783v2

    arXiv:2609.17783v2 Announce Type: replace Abstract: Streaming binary response data arise in many applications, including virtual learning platforms where student responses are collected sequentially and estimated probabilities of correct responses may inform future question assignment. Logistic regression provides an interpretable framework for such analysis, but the data may support multiple plausible predictor subsets. Bayesian model averaging (BMA) accounts for such model uncertainty by averaging inference across competing models. Since repeatedly applying BMA to accumulating data can be computationally burdensome, we develop an online implementation based on renewable estimation. At each update, retained summaries and new observations are used to approximate the joint posterior of models and model-specific coefficients without revisiting historical data, enabling point and interval estimates that incorporate model uncertainty. Simulations show that online BMA closely approximates its offline counterpart in coefficient estimation, model and variable importance, prediction, and interval estimation of predictive probabilities, while substantially reducing computational cost. In a virtual learning application, we find substantial uncertainty across competing predictor subsets and large variability in credible interval widths across questions. Our method distinguishes questions with precise estimates from those with substantial uncertainty, providing information beyond point estimates for adaptive question assignment.

    https://arxiv.org/abs/2609.17783


    Noise-adjusted turnover in estimated networks

    oai:arXiv.org:2609.20044v2

    arXiv:2609.20044v2 Announce Type: replace Abstract: Economic networks are often estimated separately over two periods, and changes in their edge sets are interpreted as structural rewiring. Since both networks are estimated, observed turnover also reflects graph-selection error. We study the two-snapshot Hamming-turnover functional under a homogeneous edge-misclassification model. With known sensitivity and specificity and conditional independence of the estimated edge indicators across periods, latent turnover admits a closed-form unbiased adjustment based only on observed turnover and the two estimated graph sizes. We then examine the effects of sparsity, calibration error and dependence across periods. When the number of true links is proportional to $p$, a false-positive probability of order $p^{-1}$ generates expected spurious turnover of order $p$. An error of the same order in calibrating the false-positive probability can likewise leave order-$p$ bias after adjustment. We also derive the bias induced by cross-period dependence and give sufficient conditions for consistency relative to network size. Numerical results illustrate the finite-sample implications.

    https://arxiv.org/abs/2609.20044


    Sampling-Based Batch Sequential Design by Stein Variational Gradient Descent

    oai:arXiv.org:2609.20583v2

    arXiv:2609.20583v2 Announce Type: replace Abstract: Many real-world experimental design problems require a batch of experimental runs across stages, in which multiple points are selected and evaluated at each stage. However, most work in the design literature is focused on fully sequential (point-by-point) methods. This paper proposes a sampling-based framework to systematically convert a fully sequential method to a batch sequential method. In particular, Stein variational gradient descent (SVGD) is adapted to efficiently sample a batch of points from a properly constructed target distribution while balancing the individual utility and the batch diversity. We address challenges that arise in using SVGD for experimental designs, including constrained design regions and near-uniform target distributions. We apply the proposed method to obtain batch versions of the state-of-the-art fully sequential methods, and demonstrate their performance through extensive numerical studies.

    https://arxiv.org/abs/2609.20583


    Trajectory Entropy Reinforcement Learning for Robust Robot Motor Skill Learning

    oai:arXiv.org:2505.04193v2

    arXiv:2505.04193v2 Announce Type: replace-cross Abstract: Simplicity is a critical inductive bias for designing data-driven controllers, especially when robustness is important. Despite the impressive results of deep reinforcement learning in complex control tasks, it is prone to capturing intricate and spurious correlations between observations and actions, leading to failure under slight perturbations to the environment. To tackle this problem, in this work we introduce a novel inductive bias towards simple policies in reinforcement learning. The simplicity inductive bias is introduced by minimizing the entropy of entire action trajectories, corresponding to the number of bits required to describe information in action trajectories after the agent observes state trajectories. Our reinforcement learning agent, Trajectory Entropy Reinforcement Learning, is optimized to minimize the trajectory entropy while maximizing rewards. We show that the trajectory entropy can be effectively estimated by learning a variational parameterized action prediction model, and use the prediction model to construct an information-regularized reward function. Furthermore, we construct a practical algorithm that enables the joint optimization of models, including the policy and the prediction model. Experimental evaluations on several high-dimensional locomotion tasks show that our learned policies produce more cyclical and consistent action trajectories, and achieve superior performance, and robustness to noise and dynamic changes than the state-of-the-art.

    https://arxiv.org/abs/2505.04193


    Fidel-TS: A High-Fidelity Multimodal Benchmark for Time Series Forecasting

    oai:arXiv.org:2509.24789v5

    arXiv:2509.24789v5 Announce Type: replace-cross Abstract: The evaluation of time series forecasting models is hindered by a lack of high-quality benchmarks, leading to overestimated assessments of progress. Existing datasets suffer from issues ranging from small-scale, low-frequency, pre-training data contamination in unimodal designs to the temporal and description leakage prevalent in early multimodal designs. To address this, we formalize the core principles of high-fidelity benchmarking, focusing on data sourcing integrity, leak-free design, and structural clarity. We introduce Fidel-TS, a new large-scale benchmark built from these principles. Our experiments reveal the limitations of prior benchmarks and the potential discrepancies in model evaluation, providing new insights into multiple existing unimodal and multimodal forecasting models and LLMs across various evaluation tasks.

    https://arxiv.org/abs/2509.24789


    Nonnegative Matrix Factorization in the Component-Wise L1 Norm for Sparse Data

    oai:arXiv.org:2603.29715v2

    arXiv:2603.29715v2 Announce Type: replace-cross Abstract: Nonnegative matrix factorization (NMF) approximates a nonnegative matrix, X, by the product of two nonnegative factors, WH, where W has r columns and H has r rows. In this paper, we consider NMF using the component-wise L1 norm as the error measure (L1-NMF), which is suited for data corrupted by heavy-tailed noise, such as Laplace noise or salt and pepper noise, or in the presence of outliers. Our first contribution is an NP-hardness proof for L1-NMF, even when r=1, in contrast to the standard NMF that uses least squares. Our second contribution is to analyze, under simplified probabilistic assumptions, how the sparsity in the data enforces zero solution in the optimal scalar update in the factors of L1-NMF when all the other entries are kept fixed. This provides an intuition of the connection between the sparsity of the L1-NMF factors with the sparsity of the input. Even though sparsity favors interpretability, if the data is affected by false zeros, too sparse solutions might degrade the model. Our third contribution is a new, more general, L1-NMF model for sparse data, dubbed weighted L1-NMF (wL1-NMF), where the sparsity of the factorization is controlled by adding a penalization parameter to the entries of WH associated with zeros in the data. The fourth contribution is a new coordinate descent (CD) approach for wL1-NMF, denoted as sparse CD (sCD), where each subproblem is solved by a weighted median algorithm. Although it lacks convergence guarantees to a stationary point, sCD is, to the best of our knowledge, the first algorithm for L1-NMF whose complexity scales with the number of nonzero entries in the data, making it efficient in handling large-scale, sparse data. We perform extensive numerical experiments on synthetic and real-world data, including imaging mass spectrometry and topic modeling, to show the effectiveness of our new proposed model (wL1-NMF) and algorithm (sCD).

    https://arxiv.org/abs/2603.29715


    Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts

    oai:arXiv.org:2609.18366v2

    arXiv:2609.18366v2 Announce Type: replace-cross Abstract: Reliable agent evaluation is complicated by automatic harness optimization, which repeatedly uses a released benchmark $B_{\mathrm{rel}}$ to guide a Proposer that edits prompts, memory, retrieval, tools, and control code around a fixed target agent. Task holdout varies semantic tasks but leaves the benchmark protocol fixed, so a "bad genius" Proposer can produce a cheating harness whose released-benchmark gain depends on a benchmark-wide shortcut. We introduce Counterfactual Harness Search and Evolution (CHASE), which casts harness evolution as constraint generation over validity-preserving benchmark counterfactuals. After each Proposer update, a Challenger searches for an executable protocol transformation with large gain destruction. A validity firewall checks that task semantics are preserved, while a confirmation set determines whether the counterfactual enters a finite archive. We formalize an exact shortcut-neutralized benchmark $B_0$ and establish statistical guarantees linking finite counterfactual archives to $B_0$ and characterizing sequential Challenger search. We evaluate CHASE on a synthetic benchmark and on OfficeQA, where CHASE retains strong released-benchmark gains while substantially reducing gain destruction under valid protocol changes.

    https://arxiv.org/abs/2609.18366


    Parallelism, critical windows, and separations among diffusion language models

    oai:arXiv.org:2609.20539v2

    arXiv:2609.20539v2 Announce Type: replace-cross Abstract: A popular selling point of diffusion large language models (dLLMs) is their capacity for parallelism: the ability to generate sequences of text far more efficiently than autoregressive models, which require one forward pass per token. Yet among the many competing paradigms for dLLMs, from masked to uniform to Gaussian diffusion, principled understanding of how these different proposals compare in parallelism remains limited. In this work, we initiate a fine-grained comparison of the capacity for parallelism among these three leading approaches and prove the following: - Uniform and Gaussian diffusion can sample in a number of forward passes which scales with the dual total correlation of the underlying distribution, a measure of intrinsic complexity which can be much smaller than the context length. Previously, it was only known how to achieve this using masked diffusion. - For a certain family of random empirical measures, we show that $\widetilde{\Theta}(\sqrt{d})$ forward passes are necessary and sufficient to sample using uniform or Gaussian diffusion, yet there exist approximate score oracles for which $\widetilde{\Omega}(d)$ forward passes are needed for masked diffusion. This establishes the first provable separation in parallelism between the three prevailing dLLM paradigms. Contrary to popular intuition that masked diffusions are harder to parallelize because they must commit to token values, the latter separation instead comes from the fact that the critical windows in masked diffusion sampling are asymptotically narrower than those in uniform and Gaussian diffusion sampling.

    https://arxiv.org/abs/2609.20539