🧠 The reasoning: text tokens are discrete, drawn from a fixed vocabulary of ~50,000 entries, each mapping to a fixed point in embedding space.
Image tokens are continuous vectors of a thousand floating point numbers, so they can pack far more information into the same space.
🎮 The test: a tower-defense game where the AI plays in JSON mode without graphics, run for hours with two context variants — the same game state as a plain text string, and as an image rendered via PixelPipe.…
mobiles Kino #Uckermark groved sich gerade eine bevor es los geht.
Got TinyUSB on my Pico giving me a test video image over UVC; The RP2040 is a bit low on RAM; it can't fit 640x256 at RGB565, and I can't find a denser deterministic image format that Linux can take as video in, Hmm, I mean an RP2350 has more RAM, but that feels like cheating on an impossible problem.
I set the alarm to 6am and we started at 7am. My wife cycled a small morning loop (like 45min) and I did a bit larger loop to test the 30 tooth Sprocket on longer climbs.
*Learnings today:*
- mosquitos can be a real motivator to push harder
- the sprocket feels good
- the #garmin incident detection works
- 1st aid kit is a good investment
- good medical coverage in …
"To test what the software could actually see, WIRED extracted the models from the camera’s files and ran them against test images and footage recovered from the device."
This means you can also take the model and evolve adversarial images that it identifies as people to disrupt image processing. Some models will only detect a limited number of targets, so a couple of high probability targets could block others from even being detected.
🖼️ Can images replace text as LLM context? A #DeepSeek paper claims one image token carries ~10 text tokens of information at close to 100% accuracy, with 59-70% cost reduction reported by PixelPipe. ThePrimeagen put it to the test. #AI
i vastly prefer all the erotica I accidentally see on Fedi pass the Bechdel test. But also be CW'd if there's an image, because not everyone on the train behind me consented to being jumpscared by boobs.
#FediMeta #JustGirlyThings
Watching the Haiku beta 6 by @… and trying to download write image test on my own laptop before the video is done.
AI-based single-shot structured-light depth reconstruction for real-time laparoscopic surgical guidance
Wayne Wonseok Rodgers, Xiangyi Le, Seonghoon Jang, Shuwen Wei, Justin Opfermann, Michael Kam, Axel Krieger, Jin U. Kang
https://arxiv.org/abs/2608.05109 https://arxiv.org/pdf/2608.05109 https://arxiv.org/html/2608.05109
arXiv:2608.05109v1 Announce Type: new
Abstract: Significance. Accurate intraoperative depth perception is important for autonomous and semi-autonomous robotic laparoscopic surgery. Conventional fringe projection profilometry can achieve millimeter-scale accuracy but often requires multi-shot acquisition, digital-micromirror-device projection, and projector-camera synchronization, complicating integration into compact laparoscopic systems.
Aim. To develop a synchronization-free, single-shot depth-sensing platform using a passive LED-illuminated binary mask and a VQ-VAE prior with a custom U-Net depth head.
Approach. A compact projection module was coupled to one channel of a dual-channel laparoscope, while the second channel imaged the fringe-illuminated target. A Zivid 3D camera acquired reference depth for 722 paired phantom images. Zivid depth maps were reprojected into the SSLE image frame for supervised training and evaluation. The VQ-VAE encoded each input into a discrete latent representation, and a latent-space U-Net predicted depth without a separate mask-prediction branch.
Results. Using a fixed train/validation/test split, the proposed model achieved an MAE of 3.70 mm, AbsRel of 0.0326, delta=1.1 accuracy of 0.962, and delta=1.1^2 accuracy of 0.970. It achieved lower MAE than the dual U-Net MaskNet DepthNet baseline and outperformed off-the-shelf monocular depth models in MAE, AbsRel, and threshold accuracy. The pipeline operated at 26.0 Hz over 301 consecutive frames on an NVIDIA A100 GPU.
Conclusions. The LED-illuminated binary-pattern platform with latent-space depth reconstruction enables synchronization-free, video-rate endoscopic depth estimation. Results demonstrate Zivid-referenced phantom reconstruction without an explicit segmentation stage, while emphasizing the importance of dataset size and SSLE-Zivid calibration accuracy.
toXiv_bot_toot
Koopman-Operator Spectral Decomposition for Nonlinear Motion Suppression in Dynamic Contrast-Enhanced MRI of the Head and Neck
Renjie He
https://arxiv.org/abs/2607.19401 https://arxiv.org/pdf/2607.19401 https://arxiv.org/html/2607.19401
arXiv:2607.19401v1 Announce Type: new
Abstract: We build a motion suppression pipeline based on Koopman operator theory, which provides a way to turn nonlinear dynamics into linear ones by looking at the data through the right set of mathematical "lenses" (called observables). We test three versions of this idea: plain DMD that works directly on pixel values, an extended version (EDMD) that adds physically motivated features like squared intensities and spatial gradients to better capture how the MRI signal and tissue motion interact, and a neural network version that tries to learn the best features automatically. A key practical contribution is time-course repetition: we tile the entire temporal series multiple times before decomposition, which does not change the underlying dynamics but gives the algorithm more data to work with, fixing a dimensionality bottleneck that otherwise prevents the extra features from helping. The full pipeline works slice by slice, dividing each image into small overlapping blocks, applying the Koopman lifting and DMD to separate slow contrast enhancement from fast motion based on their characteristic frequencies, and blending the corrected blocks back together.
toXiv_bot_toot
COBRA2026: a large-scale multicenter pelvic cone-beam computed tomography projection dataset
Adrian Thummerer, Simon Rit, Florian Kamp, Matteo Maspero, Martjin P. W. Intven, Thomas G. Bon\'e, Christopher Kurz, Guillaume Landry, Thomas Baudier, Mustafa Kadhim, Julius Arnold, Michael Rauter, Barbara Kn\"ausl, Lukas Zimmermann
https://arxiv.org/abs/2607.20037 https://arxiv.org/pdf/2607.20037 https://arxiv.org/html/2607.20037
arXiv:2607.20037v1 Announce Type: new
Abstract: The COBRA2026 dataset is a large-scale, multicenter resource of raw radiotherapy cone-beam computed tomography (CBCT) acquisitions created for the development and evaluation of conventional and learning-based reconstruction and image-correction methods. It contains data from 867 patients undergoing pelvic radiotherapy at six European centers, acquired using Elekta and Varian imaging systems. For each case, the dataset includes raw projection data, acquisition geometry, calibration and correction information, clinically reconstructed CBCT images, and corresponding planning CT images. Vendor-specific files were anonymized and converted into open formats. Planning CT images were deformably registered to the daily CBCT anatomy, and matched projections were simulated using the corresponding acquisition geometry. All cases underwent visual quality control, and cases with substantial processing or registration errors were excluded. The approximately 950 GB dataset is divided into training, validation, and test sets containing 692, 52, and 123 cases, respectively. Projection stacks and volumetric images are provided as compressed MetaImage files, with geometry and metadata supplied in XML and YAML formats. COBRA2026 supports research on full- and sparse-view reconstruction, low-dose imaging, artifact and scatter correction, motion compensation, and synthetic CT generation. The dataset is released under the CC BY-NC 4.0 license, indexed on Zenodo (doi:10.5281/zenodo.21322350), and accompanied by openly available preprocessing and baseline reconstruction code. It also forms the basis of the COBRA2026 reconstruction challenge.
toXiv_bot_toot
Post copied from Robin Evans @… but without the AI-generated image:
"""
LEAVING META
WITHOUT LEAVING YOUR LIFE 💎
After my posting about Meta over the past two days, one thing became very clear.
A lot of people would LIKE to leave Facebook, Instagram, or both. But parts of their actual lives are all tangled up in them.
So don’t make leaving a purity test. Start somewhere. Remove Facebook from your phone and use it only on your laptop, or better still, only on your desktop when you genuinely need it.
STOP POSTING TO META. STOP FEEDING THE TIMELINE.
Move conversations with real friends to your email, text, Signal, phone calls, Mastodon account, or wherever works for you, away from Meta.
TELL PEOPLE WHERE ELSE THEY CAN EASILY FIND YOU.
If you run a business, start putting your email address and/or website URL everywhere you promote your business so Facebook is not your only front door.
Download anything you actually want to keep. Then, when you’re ready, deactivate or delete what you no longer need.
Maybe you keep Messenger for Aunt Doris. Maybe you keep Instagram because your local café apparently believes websites were abolished in 2014.
Fine.
The point isn’t ideological cleanliness.
THE POINT IS REDUCING YOUR DEPENDENCE.
Every conversation you move elsewhere is one conversation Meta no longer owns the doorway to.
YOU DON'T HAVE TO ESCAPE THE META CAGE IN ONE HEROIC LEAP.
Just start opening the door.
No description
#Meta #Facebook #Instagram #SocialMedia #Fediverse #people #DEPENDENCE #Messenger
Public Aug 9, 2026 23:16 • Edited Aug 17, 2026 23:29 • EN
76 Boosts
3 Quotes
71 Favorites
"""