Where You Tap Matters: A Probe-and-Model Benchmark for Open-Set RF Fingerprinting
Gabriele Oligeri, Savio Sciancalepore, Ingrid Huso, Fatima Al-Mousawi
https://arxiv.org/abs/2607.21564 https://arxiv.org/pdf/2607.21564 https://arxiv.org/html/2607.21564
arXiv:2607.21564v1 Announce Type: new
Abstract: Radio Frequency Fingerprint Identification (RFFI) enables transmitter identification at the physical layer by learning device-specific impairments from received signals, yet the literature is inconsistent about where in the receiver chain those samples should be collected. Since distinct transformations are applied to the signal by the different receiver operations, i.e., carrier recovery, gain normalization, pulse shaping, and timing recovery, they can either tighten within-transmitter variability or suppress the features RFFI requires for classification. We present a systematic real-world evaluation of open-set, reconstruction-error RFFI using data collected at five probe points along a standard BPSK receiver chain. Our results show that RFFI is strongly probe-dependent: timing recovery and, to a lesser extent, carrier recovery enable low false-acceptance operation with limited in-distribution-out-of-distribution overlap, whereas other stages often require a false-acceptance ratio above 0.1 to achieve a true-acceptance ratio of 0.9. To test the validity of our findings across model selection, we benchmark several LLM-designed autoencoders using a controlled pipeline that holds preprocessing and MSE scoring fixed. These architectures confirm that RFFI is probe-dependent. Moreover, they do not outperform the baseline at the chosen operating point and typically increase training time. Overall, probe selection dominates reconstruction-based open-set RFFI performance, more than the autoencoder complexity.
toXiv_bot_toot
Spectral phase transitions in Gaussian multi-index models
Florent Krzakala, Pierre Mergny, Vanessa Piccolo
https://arxiv.org/abs/2608.12183 https://arxiv.org/pdf/2608.12183 https://arxiv.org/html/2608.12183
arXiv:2608.12183v1 Announce Type: new
Abstract: Recovering a low-dimensional latent subspace from nonlinear observations of Gaussian covariates in high dimensions is a fundamental problem in feature learning. Here, we consider Gaussian multi-index models in which the covariates $\boldsymbol{x}_i \stackrel{\mathrm{i.i.d.}}{\sim} \mathcal{N}(0,\boldsymbol{I}_d)$ and the responses $\boldsymbol{y}_i$ depend on $\boldsymbol{x}_i$ only through its projection onto an unknown $r$-dimensional subspace. Earlier work based on approximate message passing (AMP) identified a sharp threshold for weak recovery [Troiani et al., 2025], raising the question of whether it can be attained, without side information, by a spectral method. We answer this affirmatively and develop a general random matrix theory for matrix-valued spectral estimators of the form \[\boldsymbol{D}_n=\frac{1}{n}\sum_{i=1}^n\boldsymbol{T}(\boldsymbol{y}_i)\otimes\boldsymbol{x}_i\boldsymbol{x}_i^\top,\] where $\boldsymbol{T}$ is an arbitrary bounded symmetric matrix-valued preprocessing map of fixed dimension. As $n,d \to \infty$ with $n/d\to\alpha$, we prove that the empirical spectral measure of $\boldsymbol{D}_n$ converges almost surely to a deterministic compactly supported distribution characterized by a matrix-valued self-consistent equation. We then establish a spectral phase transition for the largest eigenvalue: below threshold it sticks to the bulk edge, while above threshold an outlier emerges. We characterize the outlier location through a finite-dimensional deterministic equation and show that the associated spectral estimator achieves weak recovery of the latent subspace. Finally, we prove that the AMP-derived preprocessing of [Defilippis et al., 2025] is optimal among all bounded matrix-valued preprocessing maps of any fixed dimension. Its transition coincides with the AMP weak-recovery threshold, proving the general spectral conjecture of [Defilippis et al., 2025].
toXiv_bot_toot
COBRA2026: a large-scale multicenter pelvic cone-beam computed tomography projection dataset
Adrian Thummerer, Simon Rit, Florian Kamp, Matteo Maspero, Martjin P. W. Intven, Thomas G. Bon\'e, Christopher Kurz, Guillaume Landry, Thomas Baudier, Mustafa Kadhim, Julius Arnold, Michael Rauter, Barbara Kn\"ausl, Lukas Zimmermann
https://arxiv.org/abs/2607.20037 https://arxiv.org/pdf/2607.20037 https://arxiv.org/html/2607.20037
arXiv:2607.20037v1 Announce Type: new
Abstract: The COBRA2026 dataset is a large-scale, multicenter resource of raw radiotherapy cone-beam computed tomography (CBCT) acquisitions created for the development and evaluation of conventional and learning-based reconstruction and image-correction methods. It contains data from 867 patients undergoing pelvic radiotherapy at six European centers, acquired using Elekta and Varian imaging systems. For each case, the dataset includes raw projection data, acquisition geometry, calibration and correction information, clinically reconstructed CBCT images, and corresponding planning CT images. Vendor-specific files were anonymized and converted into open formats. Planning CT images were deformably registered to the daily CBCT anatomy, and matched projections were simulated using the corresponding acquisition geometry. All cases underwent visual quality control, and cases with substantial processing or registration errors were excluded. The approximately 950 GB dataset is divided into training, validation, and test sets containing 692, 52, and 123 cases, respectively. Projection stacks and volumetric images are provided as compressed MetaImage files, with geometry and metadata supplied in XML and YAML formats. COBRA2026 supports research on full- and sparse-view reconstruction, low-dose imaging, artifact and scatter correction, motion compensation, and synthetic CT generation. The dataset is released under the CC BY-NC 4.0 license, indexed on Zenodo (doi:10.5281/zenodo.21322350), and accompanied by openly available preprocessing and baseline reconstruction code. It also forms the basis of the COBRA2026 reconstruction challenge.
toXiv_bot_toot