Tootfinder

Opt-in global Mastodon full text search. Join the index!

No exact results. Similar results found.
@arXiv_qbioOT_bot@mastoxiv.page
2026-07-21 07:56:08

Orientation Reading by Production Vision-Language Models on Optotype Charts: A Controlled Multi-Model Evaluation Across Reasoning Modes, Prompts, and Access Modalities
Shahryar Wasif, Avneek Sandhu, Bin Hu
arxiv.org/abs/2607.16595 arxiv.org/pdf/2607.16595 arxiv.org/html/2607.16595
arXiv:2607.16595v1 Announce Type: new
Abstract: OBJECTIVES: Vision-language models are increasingly used to interpret medical and everyday images through consumer chat interfaces, yet their ability to read orientation - the single perceptual operation tested by the tumbling-E acuity optotype - is poorly characterized on the surfaces through which they are actually used. METHODS: We evaluated four production vision-language models (referred to as Claude, GPT, GROK, and Gemini) through their consumer chat interfaces on a locked set of seven optotype charts: four uniform tumbling-E charts (one per cardinal orientation), two mixed-orientation tumbling-E charts, and one Snellen letter chart as a specificity control. Each model was run in two reasoning modes (Fast and Thinking) under two prompt variants (with and without an explicit orientation-decoding rule) by up to three operators. The corpus comprised 920 scoreable trials and 50,420 glyph judgements. The primary outcome was glyph-level accuracy against the chart's designed orientation, summarized with Wilson 95% confidence intervals. RESULTS: Accuracy ranged from 43.0% to 97.0% across models on identical charts, and the strongest model depended on reasoning mode (GPT 97.0% in Fast mode; GROK 96.6% in Thinking mode). Errors were not random but collapsed onto a model-specific attractor direction. Models were 96-100% internally self-consistent yet ranged widely in accuracy, dissociating reliability from validity. An answer-key-free ensemble-consensus estimate tracked accuracy closely (r = 0.998). For one model, consumer-interface accuracy fell 25-27 points below programmatic access, almost entirely on a single orientation. CONCLUSIONS: A single accuracy figure conceals clinically relevant, orientation-specific failure modes; vision-language models should be evaluated along multiple axes and on the deployment surface before image-interpretation outputs are trusted.
toXiv_bot_toot

@arXiv_qbioNC_bot@mastoxiv.page
2026-07-22 07:57:40

Eccentricity-Constrained CNN Training Reveals Adaptive Information Coding Around the Visual Field
Dylan M. Diaz, Margaret M. Henderson
arxiv.org/abs/2607.19316 arxiv.org/pdf/2607.19316 arxiv.org/html/2607.19316
arXiv:2607.19316v1 Announce Type: new
Abstract: In the primate visual system, center-preferring cortical populations have higher spatial resolution and overlap face- and word-selective regions while periphery-preferring populations have lower spatial resolution and overlap scene-selective regions. This "eccentricity bias" may reflect differential task-relevance: central vision may better support fine-grained tasks like face recognition and reading, while peripheral vision may better support scene understanding. To test whether eccentricity-dependent coding can emerge from natural experience, we used egocentric video and eye-tracking data from the Visual Experience Dataset (VEDB). We trained ResNet-18 models using contrastive learning (SimCLR) on frames modified to isolate different eccentricities (gaze-contingent fovea-only crops, periphery-only crops, and periphery-only crops with a NeuroFovea transform applied). We evaluated downstream task performance and model alignment with human fMRI data (Natural Scenes Dataset; encoding models). In-domain VEDB frame classification showed systematic differences between fovea- and periphery-only models across categories, indicating differential informativeness across tasks. On downstream classification, VEDB-pretrained models generalized better to scene categorization (Places365) than face recognition (VGGFace2), with fovea-only models stronger on both. Across visual cortex, VEDB-pretrained models matched neural predictivity of models trained on mid-sized non-egocentric datasets (ImageNet-100), suggesting egocentric data supports emergence of cortically-aligned representations. In scene-selective cortex (PPA, RSC), periphery-only models held a small but consistent advantage in explained variance over fovea-only models, suggesting these regions are aligned with peripheral statistics. Together, these results suggest egocentric experience may adaptively constrain cortical information processing.
toXiv_bot_toot

@arXiv_csGR_bot@mastoxiv.page
2026-07-21 07:34:37

Feature-Guided Diffusion for Non-Differentiable Inverse Rendering
Andrei-Timotei Ardelean, Michael Fischer, Tim Weyrich, Tom\'a\v{s} Iser
arxiv.org/abs/2607.17411 arxiv.org/pdf/2607.17411 arxiv.org/html/2607.17411
arXiv:2607.17411v1 Announce Type: new
Abstract: Inverse rendering is traditionally solved via differentiable renderers and gradient descent, which requires substantial problem-specific engineering and is prone to getting stuck in local minima due to ambiguities. Derivative-free approaches alleviate engineering requirements, but often heavily depend on a good problem initialization. In this work, we propose Feature-Informed Diffusion Evolution (FIDE), a fully black-box framework that requires no gradients or specific initialization: the renderer is treated as an opaque function whose only requirement is to produce images. Our key insight is feature guiding: rather than reducing each candidate rendering to a scalar loss value, we use a Vision Transformer (ViT) to extract dense visual features from it. We subsequently use these features to train a diffusion-based candidate proposal model, allowing the network to use visual cues to predict parameters that would match the target image. The candidate solutions proposed by this diffusion model are then refined in a closed loop with a CMA evolution strategy, continuously narrowing the proposal region as optimization progresses. We validate across diverse inverse problems from path tracing, vector splines, Voronoi shaders, and robotics, and demonstrate that feature-guiding substantially improves convergence over scalar-loss baselines and reliably escapes local minima where gradient-based methods stall.
toXiv_bot_toot

@Techmeme@techhub.social
2026-08-10 10:15:03

Meta releases Muse Glimmer, a new open-weight model, and will release an open-weight version of its most advanced model, Muse Spark 1.2, in the coming weeks (Meghan Bobrowsky/Wall Street Journal)
wsj.com/tech/ai/mar…

@arXiv_qbioNC_bot@mastoxiv.page
2026-07-21 09:41:50

Replaced article(s) found for q-bio.NC. arxiv.org/list/q-bio.NC/new
[1/1]:
- The Illusion-Illusion: Vision Language Models See Illusions Where There Are None
Tomer Ullman
arxiv.org/abs/2412.18613 mastoxiv.page/@arXiv_qbioNC_bo
- An Intelligent Infrastructure as a Foundation for Modern Science
Satrajit S. Ghosh
arxiv.org/abs/2508.10051 mastoxiv.page/@arXiv_qbioNC_bo
- The embodied brain: Bridging the brain, body, and behavior with biorealistic neuromechanical models
Sibo Wang-Chen, Pavan Ramdya
arxiv.org/abs/2601.08056 mastoxiv.page/@arXiv_qbioNC_bo
- Microsecond-precision sound localization emerges from slow equilibrium dynamics
Toshio Irino
arxiv.org/abs/2607.03890 mastoxiv.page/@arXiv_qbioNC_bo
- A portable solution for simultaneous human movement and mobile EEG acquisition: readiness potenti...
Contreras-Altamirano, Klapprott, Jacobsen, Maanen, Welzel, Debener
arxiv.org/abs/2501.05378 mastoxiv.page/@arXiv_csNE_bot/
- Geometric origin of adversarial vulnerability in deep learning
Yixiong Ren, Wenkang Du, Jianhui Zhou, Haiping Huang
arxiv.org/abs/2509.01235 mastoxiv.page/@arXiv_csLG_bot/
- Universal Approximation Theorems for Dynamical Systems with Infinite-Time Horizon Guarantees
Abel Sagodi, Il Memming Park
arxiv.org/abs/2602.08640 mastoxiv.page/@arXiv_mathDS_bo
toXiv_bot_toot

@arXiv_csHC_bot@mastoxiv.page
2026-08-12 08:28:14

The Signal Rail: A Deterministic Motion Grammar for Communicating Conversational Agent State in Terminal Interfaces
Matteo Grella
arxiv.org/abs/2608.10689 arxiv.org/pdf/2608.10689 arxiv.org/html/2608.10689
arXiv:2608.10689v1 Announce Type: new
Abstract: Terminal interfaces to conversational agents report rich internal state (listening, thinking, executing tools, awaiting input, failing) almost entirely through text, while the motion channel beside it, the one peripheral vision monitors without reading, carries a single bit: alive. We present the Signal Rail, a one-row terminal status instrument that gives that channel a grammar. Four ideas govern it: spatial semantics (input, processing, and output zones, with direction as meaning), a motion grammar (one kinetic rule per state, never color alone), determinism (frames as a pure function of explicit inputs, golden-frame testable), and honesty (no invented progress or activity). We contribute a 45-section normative specification and a reference implementation inside a working full-duplex local voice agent driven by real signals.
toXiv_bot_toot

@Techmeme@techhub.social
2026-06-28 05:40:50

Masayoshi Son questioned Musk's orbital AI data centers, noting electricity is just 7% of costs and the AI race will be won on Earth within a few years (Tim Higgins/Wall Street Journal)
wsj.com/tech/why-one-of-techs-

@arXiv_csHC_bot@mastoxiv.page
2026-08-12 08:23:23

Automatic Field-of-View Adjustment for a View-Expansive Microscope via LSTM-Based Gaze and Pipette Motion Interpretation
Kenta Yokoe, Takuya Hara, Tadayoshi Aoyama
arxiv.org/abs/2608.10401 arxiv.org/pdf/2608.10401 arxiv.org/html/2608.10401
arXiv:2608.10401v1 Announce Type: new
Abstract: Intracytoplasmic sperm injection (ICSI) operators frequently adjust the field-of-view (FOV) during procedures, which interrupts workflow and increases procedure time. Conventional microscopes require manual objective lens switching and illumination adjustments to achieve different FOV sizes. We propose an AI-based automatic FOV adjustment method integrated with a view-expansive microscope. This microscope enables the simultaneous acquisition of a large FOV and high-resolution images using a single objective lens through multiview imaging with galvanometer mirrors and high-speed vision, thereby eliminating the need for physical lens exchanges. Our method utilizes a long short-term memory (LSTM) model to predict the appropriate FOV size based on real-time analysis of the pipette's position and velocity, combined with the operator's gaze position. The AI model is trained using ICSI procedure data from an expert with over five years of micromanipulation experience. Experimental evaluation with novice operators reveals that the proposed automatic FOV adjustment system significantly improves the ICSI procedure speed, reducing the average task completion time from 60.5 to 48.0 s (p < 0.001). The experiments also demonstrate that this improvement enables novice operators to achieve ICSI working speeds equivalent to those of expert operators.
toXiv_bot_toot

@arXiv_csHC_bot@mastoxiv.page
2026-08-12 09:04:18

Replaced article(s) found for cs.HC. arxiv.org/list/cs.HC/new
[1/2]:
- Music Interpretation and Emotion Perception: A Computational and Neurophysiological Investigation
Lyberatos, Kantarelis, Zioga, Anagnostopoulou, Stamou, Georgaki
arxiv.org/abs/2506.01982 mastoxiv.page/@arXiv_csHC_bot/
- Selective Prediction Reduces the Negative Effects of Automation Bias Overall but Increases False ...
Jabbour, Fouhey, Banovic, Shepard, Kazerooni, Sjoding, Wiens
arxiv.org/abs/2508.07617 mastoxiv.page/@arXiv_csHC_bot/
- Sighted by Default: Addressing Implicit Vision Assumptions in Real-Time VLM Assistance for BLV Users
Yi Zhao, Siqi Wang, Qiqun Geng, Erxin Yu, Jing Li
arxiv.org/abs/2511.00945 mastoxiv.page/@arXiv_csHC_bot/
- UXCascade: Scalable Usability Testing with Simulated User Agents
Steffen Holter, Eunyee Koh, Mustafa Doga Dogan, Gromit Yeuk-Yin Chan
arxiv.org/abs/2601.15777 mastoxiv.page/@arXiv_csHC_bot/
- Does Explanation Correctness Matter? Linking Computational XAI Evaluation to Human Understanding
Gregor Baer, Chao Zhang, Isel Grau, Pieter Van Gorp
arxiv.org/abs/2603.25251 mastoxiv.page/@arXiv_csHC_bot/
- Beyond Screenshots: Evaluating VLMs' Understanding of UI Animations
Chen Liang, Xirui Jiang, Naihao Deng, Eytan Adar, Anhong Guo
arxiv.org/abs/2604.26148 mastoxiv.page/@arXiv_csHC_bot/
- Quieting the Cobwebs: Browser Interaction for Visual Floaters
Kenneth Ge, Jinglin Li, Shikhar Ahuja
arxiv.org/abs/2605.12739 mastoxiv.page/@arXiv_csHC_bot/
- Patients With Personality: Realistic Patient Simulation through Controlled Diversity and Selectiv...
Moritz Schlager, et al.
arxiv.org/abs/2606.17441 mastoxiv.page/@arXiv_csHC_bot/
- TailVis: Expressive Chart Refinement Preserving Data-Binding Integrity
Yumin Song, Seokhyeon Park, Soohyun Lee, Aeri Cho, Hyeon Jeon, John Joon Young Chung, Jinwook Seo
arxiv.org/abs/2607.25386 mastoxiv.page/@arXiv_csHC_bot/
- How Usable Are Geospatial Foundation Models? A Systematic Evaluation of 89 Models
Robin Young, Artyom Gabtraupov, Kenzy Soror, Srinivasan Keshav
arxiv.org/abs/2608.03804 mastoxiv.page/@arXiv_csHC_bot/
toXiv_bot_toot

@arXiv_csHC_bot@mastoxiv.page
2026-08-12 08:22:32

Visual-to-Haptic Augmentation in XR: A Wearable Glove for Perceptual Grounding in Multimodal Interaction
Faisal Mohd, Hamdi Elsaddik, Erhan Baturay Onural, Jihong Zhang, Fedwa Laamarti, Abdulmotaleb El Saddik
arxiv.org/abs/2608.10368 arxiv.org/pdf/2608.10368 arxiv.org/html/2608.10368
arXiv:2608.10368v1 Announce Type: new
Abstract: Extended Reality (XR) systems increasingly deliver high-fidelity visual and auditory experiences, yet tactile perception remains comparatively underutilized as a modality for enriching embodied interaction. This work presents a visual-to-haptic wearable glove and a feature-based visual-to-haptic mapping algorithm that translates spatial and temporal visual features from images and videos into distributed vibrotactile patterns. The proposed method extracts motion, edge, and brightness cues and fuses them into actuator-level intensity maps aligned with a 29-actuator glove arranged in a five-by-seven layout.
The system is implemented through a modular four-layer architecture comprising the XR environment, media content handling, visual-to-haptic processing, and embedded haptic hardware. A within-subject user study (N = 20) compared visual-only interaction with visual-plus-haptic augmentation across texture-based and dynamic video scenarios. Results indicate that tactile augmentation significantly improves perceived realism in dynamic video scenarios and enhances immersion and visual-tactile correspondence across conditions, with stronger and more consistent effects observed for dynamic visual events.
While the current implementation operates in a single-user, offline-synchronized configuration, the findings demonstrate that vision-driven tactile augmentation can function as a perceptual enhancement layer within multimodal XR systems. Such a layer may provide a foundation for future socially enriched XR environments where coherent multisensory grounding supports higher-level interaction and communication.
toXiv_bot_toot