I set the alarm to 6am and we started at 7am. My wife cycled a small morning loop (like 45min) and I did a bit larger loop to test the 30 tooth Sprocket on longer climbs.
*Learnings today:*
- mosquitos can be a real motivator to push harder
- the sprocket feels good
- the #garmin incident detection works
- 1st aid kit is a good investment
- good medical coverage in …
In other news, my first 15x11" (roughly DIN A3) print came out with only minor issues, none of them show stoppers and I know what needs to be done to improve. So this is just a test print and it's a bit hard to capture, but the overall presence & level of detail is so amazing at this size! Hard to go back to smaller sizes now :)
To get here, I've been doing countless tests of different recipes and hardware (e.g. I had to disassemble my UV light to create a more uniform/unfocused li…
I can't get over how bad the docker ecosystem is.
"we provide a docker image, just use that!"
*install container image & run it - it doesn't work*
"oops, we didn't test it"
🧠 The reasoning: text tokens are discrete, drawn from a fixed vocabulary of ~50,000 entries, each mapping to a fixed point in embedding space.
Image tokens are continuous vectors of a thousand floating point numbers, so they can pack far more information into the same space.
🎮 The test: a tower-defense game where the AI plays in JSON mode without graphics, run for hours with two context variants — the same game state as a plain text string, and as an image rendered via PixelPipe.…
i vastly prefer all the erotica I accidentally see on Fedi pass the Bechdel test. But also be CW'd if there's an image, because not everyone on the train behind me consented to being jumpscared by boobs.
#FediMeta #JustGirlyThings
🖼️ Can images replace text as LLM context? A #DeepSeek paper claims one image token carries ~10 text tokens of information at close to 100% accuracy, with 59-70% cost reduction reported by PixelPipe. ThePrimeagen put it to the test. #AI
Koopman-Operator Spectral Decomposition for Nonlinear Motion Suppression in Dynamic Contrast-Enhanced MRI of the Head and Neck
Renjie He
https://arxiv.org/abs/2607.19401 https://arxiv.org/pdf/2607.19401 https://arxiv.org/html/2607.19401
arXiv:2607.19401v1 Announce Type: new
Abstract: We build a motion suppression pipeline based on Koopman operator theory, which provides a way to turn nonlinear dynamics into linear ones by looking at the data through the right set of mathematical "lenses" (called observables). We test three versions of this idea: plain DMD that works directly on pixel values, an extended version (EDMD) that adds physically motivated features like squared intensities and spatial gradients to better capture how the MRI signal and tissue motion interact, and a neural network version that tries to learn the best features automatically. A key practical contribution is time-course repetition: we tile the entire temporal series multiple times before decomposition, which does not change the underlying dynamics but gives the algorithm more data to work with, fixing a dimensionality bottleneck that otherwise prevents the extra features from helping. The full pipeline works slice by slice, dividing each image into small overlapping blocks, applying the Koopman lifting and DMD to separate slow contrast enhancement from fast motion based on their characteristic frequencies, and blending the corrected blocks back together.
toXiv_bot_toot
Replaced article(s) found for eess.AS. https://arxiv.org/list/eess.AS/new
[1/1]:
- Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
Muhammad Shakeel, Yui Sudo, Yifan Peng, Chyi-Jiunn Lin, Shinji Watanabe
https://arxiv.org/abs/2508.20474 https://mastoxiv.page/@arXiv_eessAS_bot/115110974009150613
- CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR
Muhammad Shakeel, Yosuke Fukumoto, Chikara Maeda, Chyi-Jiunn Lin, Shinji Watanabe
https://arxiv.org/abs/2601.22792 https://mastoxiv.page/@arXiv_eessAS_bot/116000207024295325
- How Much Does Machine Identity Matter in Anomalous Sound Detection at Test Time?
Kevin Wilkinghoff, Keisuke Imoto, Zheng-Hua Tan
https://arxiv.org/abs/2602.16253 https://mastoxiv.page/@arXiv_eessAS_bot/116096185732811365
- LMU-Based Sequential Learning and Posterior Ensemble Fusion for Cross-Domain Infant Cry Classific...
Niloofar Jazaeri, Hilmi R. Dajani, Marco Janeczek, Martin Bouchard
https://arxiv.org/abs/2603.02245 https://mastoxiv.page/@arXiv_eessAS_bot/116169771215037748
- Adapting a Text-to-Audio Model for Room Impulse Response Generation
Kirak Kim, Sungyoung Kim
https://arxiv.org/abs/2603.09708 https://mastoxiv.page/@arXiv_eessAS_bot/116209762413602825
- Repurposing Image Diffusion Models for Training-Free Music Style Transfer on Mel-spectrograms
Heehwan Wang, Joonwoo Kwon, Sooyoung Kim, Jungwoo Seo, Shinjae Yoo, Yuewei Lin, Jiook Cha
https://arxiv.org/abs/2411.15913 https://mastoxiv.page/@arXiv_csSD_bot/113548024475383386
- DeePen: Penetration Testing for Audio Deepfake Detection
M\"uller, Kawa, Stan, Doan, Jung, Choong, Sperl, B\"ottinger
https://arxiv.org/abs/2502.20427 https://mastoxiv.page/@arXiv_csCR_bot/114097333876265997
- Re-evaluating Minimum Bayes Risk Decoding for Automatic Speech Recognition
Yuu Jinnai
https://arxiv.org/abs/2510.19471 https://mastoxiv.page/@arXiv_csCL_bot/115422969877240889
- Aliasing-Free Neural Audio Synthesis
Yicheng Gu, Junan Zhang, Chaoren Wang, Jerry Li, Zhizheng Wu, Lauri Juvela
https://arxiv.org/abs/2512.20211 https://mastoxiv.page/@arXiv_csSD_bot/115773521971327576
- TiCo: Time-Controllable Spoken Dialogue Model
Kai-Wei Chang, Wei-Chih Chen, En-Pei Hu, Hung-yi Lee, James Glass
https://arxiv.org/abs/2603.22267 https://mastoxiv.page/@arXiv_csCL_bot/116283643505371784
toXiv_bot_toot
COBRA2026: a large-scale multicenter pelvic cone-beam computed tomography projection dataset
Adrian Thummerer, Simon Rit, Florian Kamp, Matteo Maspero, Martjin P. W. Intven, Thomas G. Bon\'e, Christopher Kurz, Guillaume Landry, Thomas Baudier, Mustafa Kadhim, Julius Arnold, Michael Rauter, Barbara Kn\"ausl, Lukas Zimmermann
https://arxiv.org/abs/2607.20037 https://arxiv.org/pdf/2607.20037 https://arxiv.org/html/2607.20037
arXiv:2607.20037v1 Announce Type: new
Abstract: The COBRA2026 dataset is a large-scale, multicenter resource of raw radiotherapy cone-beam computed tomography (CBCT) acquisitions created for the development and evaluation of conventional and learning-based reconstruction and image-correction methods. It contains data from 867 patients undergoing pelvic radiotherapy at six European centers, acquired using Elekta and Varian imaging systems. For each case, the dataset includes raw projection data, acquisition geometry, calibration and correction information, clinically reconstructed CBCT images, and corresponding planning CT images. Vendor-specific files were anonymized and converted into open formats. Planning CT images were deformably registered to the daily CBCT anatomy, and matched projections were simulated using the corresponding acquisition geometry. All cases underwent visual quality control, and cases with substantial processing or registration errors were excluded. The approximately 950 GB dataset is divided into training, validation, and test sets containing 692, 52, and 123 cases, respectively. Projection stacks and volumetric images are provided as compressed MetaImage files, with geometry and metadata supplied in XML and YAML formats. COBRA2026 supports research on full- and sparse-view reconstruction, low-dose imaging, artifact and scatter correction, motion compensation, and synthetic CT generation. The dataset is released under the CC BY-NC 4.0 license, indexed on Zenodo (doi:10.5281/zenodo.21322350), and accompanied by openly available preprocessing and baseline reconstruction code. It also forms the basis of the COBRA2026 reconstruction challenge.
toXiv_bot_toot