Non-Blurry Figures 🧿
非-模糊的形象 🧿
📷 Yashica 635
🎞️ Ilford FP4 Plus 125 (FF), expired 1994
If you like my work, Support by buying me a coffee or a roll of film from
PayPal https://www.paypal.com/paypalme/ydcdingsite
Wise
RTP-LLM: High-Performance Alibaba LLM Inference Engine
Boyu Tan, Jiarui Guo, Zongwei Lv, Hanbo Sun, Tong Yang, Kan Liu, Xinfei Shi, Zetao Hu, Yaxin Yu, Chi Zhang, Jianning Zhang, Xi Yang, Wei Zhang, Bo Cai, Silu Zhou, Xiyu Wang, Na He, Yinghao Yu, Wending Bao, Guiyang Huang, Yuxing Yuan, Juncheng Yin, Nan Wang, Lin Yang, Zechao Zhang, Lu Chen, Guoding Li, Tao Lan, Lin Qu
https://arxiv.org/abs/2605.29639 https://arxiv.org/pdf/2605.29639 https://arxiv.org/html/2605.29639
arXiv:2605.29639v1 Announce Type: new
Abstract: Large Language Models (LLMs) have revolutionized AI applications, but deploying them at scale presents significant challenges. We present RTP-LLM, a high-performance inference engine for industrial-scale LLM deployment, successfully deployed across Alibaba Group serving over 100 million users. RTP-LLM addresses fundamental bottlenecks through integrated design. It optimizes model loading via file-order-driven I/O and parallel I/O-communication overlapping. The Prefill-Decode Disaggregation architecture decouples compute-intensive prefill from memory-bound decode phases, combined with hierarchical multi-tiered KV cache management enabling efficient cache reuse. In addition, RTP-LLM incorporates modular speculative decoding supporting multiple algorithms, adaptive KV cache quantization, and decoupled multimodal processing, with support for multi-level parallelism.
Comprehensive evaluations across diverse model architectures (8B-235B parameters) have been conducted, where both controlled benchmarks and real production workloads are used. The results demonstrate RTP-LLM's superior performance against vLLM and SGLang: 4.7x-6.3x model loading speedup, 35-37% TTFT P95 latency reduction with 215% cache reuse improvement in production traffic scheduling, 1.12x-2.48x and 1.86x-2.52x throughput improvements in speculative decoding and multimodal inference, respectively, and 35-40% batch latency reduction with 1.9x-3.0x TTFT improvement in quantized inference. RTP-LLM's production-proven architecture and open-source availability make it a comprehensive solution for industrial LLM deployment.
toXiv_bot_toot
Perfection Because It Doesn’t Exist 🎐
完美因为完美不存在 🎐
📷 Nikon F4E
🎞️ Kentmere 400
If you like my work, Support by buying me a coffee or a roll of film from
PayPal https://www.paypal.com/paypalme/ydcdingsite
Wise
This is the most detailed picture of a human cell ever made 🧪
https://www.instagram.com/reel/DX1CGIZMKJs/?igsh=NTc4MTIwNjQ2YQ
City Silhouettes V🏙️
城市轮廓线 V 🏙️
📷 Pentax 6x7
🎞️ Kentmere 400 (6x7)
If you like my work, Support by buying me a coffee or a roll of film from PayPal https://www.paypal.com/paypalme/ydcdingsite
Since it was relevant to a discussion I just had on here and is something most people probably haven't thought about much (unless you've taken one of a handful of philosophy classes), I thought I'd try to lay out a key piece of Descartes' Meditations (#philosophy
From Licensing to Open Access: Designing a Sustainable Transition in Operational Weather Data
Emma Pidduck, Umberto Modigliani, Victoria L. Bennett, Fabio Venuti, Florian Pappenberger, Florence Rabier
https://arxiv.org/abs/2605.21673 https://arxiv.org/pdf/2605.21673 https://arxiv.org/html/2605.21673
arXiv:2605.21673v1 Announce Type: new
Abstract: This translational article documents the European Centre for Medium-Range Weather Forecasts (ECMWF) transition from a restricted data licensing model to open access under CC BY 4.0, completed in October 2025. The policy context included EU open data requirements and alignment with international data exchange frameworks. The transition was implemented through a tiered service model that kept core forecast data open while offering operationally supported delivery as a cost-recovered service. Between 2020 and 2025, ECMWF executed an iterative planning cycle: setting an annual target for revenue reduction, specifying additions to the open tier under that target, provisioning infrastructure, and assessing outcomes to update assumptions. Drawing on internal administrative records (2014 - 2025), we describe design choices, operational constraints, and early outcomes. In the six months following the end of the transition, more than 93% of previously paying organisations retained a Service Agreement, while open endpoint download volumes increased substantially. We discuss trade-offs in defining the open tier (resolution, parameters, schedule), the reduction of compliance overheads formerly associated with redistribution restrictions, and the scalability implications of global distribution. We note an emerging sustainability question as AI-based forecast products become freely available. The early evidence is consistent with the view that a tiered service model can be designed to reconcile open-access obligations with operational sustainability, subject to monitoring over longer contract renewal cycles (typically annual).
toXiv_bot_toot
Urban Illusions and Fallacies - IV 🏙️
城市的幻影和谬误 IV 🏙️
📷 Pentax 6x7
🎞️ Kentmere Pan 200 (6x7)
If you like my work, Support by buying me a coffee or a roll of film from PayPal https://paypal.com/paypalme/ydcdingsite
Same Place 🐲
同一個地方 🐲
📷 Nikon F4E
🎞️ Kentmere 400
If you like my work, Support by buying me a coffee or a roll of film from
PayPal https://www.paypal.com/paypalme/ydcdingsite
Wise