18天前更新HISTORY近 30 天历史柱高表示当天去重热搜数量
07/26—08/24 有历史数据
- 01Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and EditingarXiv:2608.02711v1 Announce Type: new Abstract: Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimodal data, particularly the lack of large-scale and geometrically consistent editing data. To address this limitation, we propose Hunyuan3D-Buffalo 1.0, a unified framework supporting 3D understanding, text-to-3D generation, instructioJunliang Ye, Kenkun Liu, Guocun Wang, Yang Li, Yansong Qu, Chunshi Wang, Jingwei Xu, Yunhan Yang, Zibo Zhao, Jiachen Xu, Jiaao Yu, Lifu Wang, Zhihao Liang, Zhuo Chen, Chunchao Guo
- 02Quo Vadis, World Modeling?arXiv:2608.02713v1 Announce Type: new Abstract: Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real-environment interaction is costly, slow, unsafe, and hard to parallelize. World modeling offers a natural intermediate proxy that allows agents to query lower-cost, more controllable feedback before committing to real actions. Classical world models instantiate this proxy primarily through future physical-state prediction, a formulation usefYu Yang, Xuemeng Yang, Licheng Wen, Lingdong Kong, Xiaobin Hu, Dongyue Lu, Wei Chow, Xiyan Huang, Yuxiang Feng, Yue Liao, Jianbiao Mei, Daocheng Fu, Rong Wu, Pinlong Cai, Ran Yi, Ying Tai, Jiangning Zhang, Botian Shi, Yong Liu, Shuicheng Yan
- 03Oh Deer, How Should I Handle This? Seasonal Priors for Selective Wildlife Annotation and ClassificationarXiv:2608.02762v1 Announce Type: new Abstract: Fine-grained wildlife classification in aerial imagery is limited not only by model performance, but also by unreliable labels: animals occupy few pixels, key visual cues vary seasonally, and modality-specific evidence can be ambiguous. We study adult-male identification in red deer, where the antler cycle defines predictable windows of reliable evidence for both annotation and prediction. Using 7,295 RGB-only, thermal-only, and matched RGB+thermalHugo Markoff, Christoph Praschl, Anton Hjalte J{\o}rgensen, Christian Emil Mogensen, Mathias Bech Skadhauge, Sara Beery, Michael {\O}rsted, David C. Schedl
- 04Confident but Unreliable: A Behavioral Safety Audit of Vision-Language Models on Brain MRIarXiv:2608.02790v1 Announce Type: new Abstract: Vision-language models (VLMs), including medical specialists, are increasingly proposed for medical imaging, yet their stated confidence is rarely evaluated separately from correctness. We use brain MRI as a controlled, high-stakes testbed for a broader failure mode in frontier multimodal systems: models can appear competent while lacking reliable self-knowledge. We present an automatically graded behavioral audit and pilot study of six instructionAmir Sabbaghziarani, Mohammadsajad Abavisani, Sergey Plis
- 05Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based SegmentationarXiv:2608.02791v1 Announce Type: new Abstract: MLLM-based segmentation faces a core segmentation trilemma: high segmentation performance, preserved dialogue ability, and fast inference. Embedding-prediction methods may disrupt language modeling through pixel-level objectives, whereas next-token generation is inefficient for dense masks. We propose All-Mask Prediction, decoupling autoregressive dialogue from non-autoregressive mask prediction. Its binary instantiation, STAMP (Simultaneous TextuaJiazhen Liu, Mingkuan Feng, Long Chen
- 06PixelUp: Zero-Shot Semantic Feature Upsampling for Fine-Grained Vision TasksarXiv:2608.02792v1 Announce Type: new Abstract: Self-supervised Vision Foundation Models (VFMs) have become essential backbones for downstream tasks due to their strong and transferable visual representations. However, their patch-token-level features are often too coarse for dense prediction tasks such as semantic segmentation and depth estimation when accurate fine-grained predictions are required. Feature upsampling methods have been developed to recover pixel-level detail but still face limiDeepank Singh, Anurag Nihal, Vedhus Hoskere
- 07SAGE: Semantic Explainability of Attention-Based Survival Models in Computational PathologyarXiv:2608.02803v1 Announce Type: new Abstract: Attention-based multiple instance learning (ABMIL) is the predominant approach for slide-level prediction in computational pathology, yet its attention maps provide only local explanations: they indicate where a model focuses but not which histological features drive its predictions or how the model behaves across a patient cohort. We present Semantic Attention Global Explanations (SAGE), a post-hoc framework that extracts global, language-groundedAbdallah Lamane, Abdul Rahman Diab, Ren-Chin Wu, William Lotter
- 08A Unified 2D Framework for DeepLesion Detection, Segmentation and Short Report GenerationarXiv:2608.02805v1 Announce Type: new Abstract: In previous work, we integrated large language models (LLMs) into the lesion segmentation model based on the ULS23 DeepLesion dataset, using short-form findings from the reports. In this study, we developed a unified 2D lesion analysis framework that integrates LLM-based reasoning, lesion bounding box detection, segmentation, and radiology report generation from the original DeepLesion dataset. In the testing phase, we achieved relatively high lesiRuida Cheng, Tejas S. Mathai, Benjamin Hou, Qingqing Zhu, Zhiyong Lu, Matthew McAuliffe, Ronald M. Summers
- 09In-Context Collapse in Vision-Language Models and How to Mitigate it?arXiv:2608.02830v1 Announce Type: new Abstract: Many-shot in-context learning (ICL) lets vision-language models (VLMs) adapt from image--label demonstrations without weight updates, and is widely assumed to improve as more demonstrations are supplied. We show the opposite: as demonstrations accumulate, a subset of VLMs undergo an \emph{in-context collapse}, a sharp, sometimes catastrophic accuracy drop spanning synthetic classification, natural-image classification, and VQA benchmarks, in some mMohammad Rostami
- 10CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded ReasoningarXiv:2608.02833v1 Announce Type: new Abstract: Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension with logical reasoning, yet current models struggle with accurate visual grounding and coherent reasoning chains. While extrinsic chain-of-thought prompting and visual cues significantly improve performance, current MLLMs lack intrinsic visual grounded reasoning capabilities, leading to inaccurate perception and reasoning disconnectedXuehang Guo, Pingyue Zhang, Ruiyi Zhang, Zhenhailong Wang, Hanrui Lyu, Heng Ji, Tong Sun, Qingyun Wang, Manling Li
- 11A Human-in-the-Loop Deep Learning Framework for Color Reconstruction of Lenticular FilmsarXiv:2608.02835v1 Announce Type: new Abstract: Historical lenticular films, such as those created with the Kodacolor process, encode color information in a distinctive spatial format. This structure requires specialized techniques for accurate color reconstruction. While recent signal processing approaches like doLCE and deep learning methods like deep-doLCE have advanced automated color recovery, they often fail with cases such as curved lenticules, low-contrast, or badly captured regions. WeSaptarshi Neil Sinha, Tiago Kleist, Giorgio Trumpy
- 12Localize, Don't Beautify: Client-Side Control of Image-Editing APIs for Cosmetic Surgery PreviewsarXiv:2608.02841v1 Announce Type: new Abstract: Ask a commercial image editor to preview a cosmetic procedure and it will often change more of the face than the request names: a nose edit can also smooth skin or alter lighting. Existing methods for confining an edit to one region require access to the model's internals, which a public editing API does not expose. We ask how much control is possible from the client side alone. In a pilot benchmark, six commercial editing configurations and one maSukhrobbek Ilyosbekov
- 13Test Time Adaptation Methods for Point Cloud Registration in Laparoscopic SurgeryarXiv:2608.02883v1 Announce Type: new Abstract: 3D point cloud registration in laparoscopic surgery estimates the transformation between an intraoperative organ reconstructed from video and its preoperative mesh. Because ground-truth transformations are unavailable for real data, supervised networks are trained on synthetic organ pairs. At test time, real reconstructions differ from synthetic data and are noisy, sparse, and occluded, which degrades correspondence estimation. Test-time adaptationNina Bodelot, Soufiane Belharbi, Eric Granger
- 14Modeling Scientific Experiment Scenes: Dataset and ModelarXiv:2608.02892v1 Announce Type: new Abstract: Scene Graph Generation (SGG) is fundamental to structured visual understanding, yet existing benchmarks focus mainly on daily life images and overlook scientific experiment scenes with specialized instruments, task-specific experimental semantics, and dense, fine-grained physical relations. These scenes are increasingly important for automated experimental analysis and smart education. To bridge this gap, we introduce PhysScene, the first SGG datasMinghao Zou, Qingtian Zeng, Shangkun Liu, Cong Liu, Paul L. Rosin, Guanghui Yue, Jun Liu, Wei Zhou
- 15RealWeather: Realistic and Scene-Faithful Weather Translation with Driving World ModelsarXiv:2608.02953v1 Announce Type: new Abstract: Realistic weather translation is valuable for developing and evaluating autonomous driving systems, yet collecting paired videos of the same scenes under different weather conditions at scale is impractical. Existing methods therefore rely on synthetic data, 3D weather editing, or geometry-conditioned generation, often compromising weather realism or scene fidelity. We propose RealWeather, a driving world model for both realistic and scene-faithfulYuwei Ning, Liangzhi Wang, Yi Xiao, Zhenhua Wu, Yun Pang, Mingkun Chan, Jichang Li, Guanbin Li
- 16Material-Segmented Per-Pixel Emissivity Correction for Thermographic Anomaly Detection in Cultural Heritage Digital TwinsarXiv:2608.02964v1 Announce Type: new Abstract: Quantitative longwave thermography of heritage surfaces is limited by the global-constant emissivity assumption in inverse-Planck temperature retrieval; on heterogeneous surfaces emissivity varies within one field of view, producing apparent-temperature artifacts that mimic and mask subsurface anomalies. We present a training-free pipeline that derives per-pixel emissivity by applying SAM 3.1 open-vocabulary segmentation to a colocated, co-calibratJonathan Klingspon, Scott McAvoy, Maurizio Seracini, Falko Kuester
- 17Qwen-3D: A Generalist 3D Vision-Language Model for Spatial UnderstandingarXiv:2608.02980v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) have achieved remarkable success on images and short videos, yet scaling them to long videos remains challenging due to frame-centric tokenization and limited context windows. 3D geometry provides a natural compression mechanism for visual streams: depth and camera pose enable observations from multiple views and time steps to be fused into a persistent, world-aligned representation. While recent 3D LMMs leverage geomLucy Lin, Ayush Jain, Yifan Liu, Katerina Fragkiadaki
- 18V-FIND: Revealing the Intrinsic Forgery Knowledge Encoded in Video Forgery DetectorsarXiv:2608.03008v1 Announce Type: new Abstract: As generated videos become increasingly realistic, reliable video forgery detection is increasingly important. Existing studies typically optimize and use video forgery detectors as black boxes, while the latent forgery-discriminative knowledge inside them remains largely unexplored. Instead of continuing to rely on resource-intensive full-model retraining to steadily improve detection performance, we ask whether video forgery detection can also beShichao Kan, Chengpeng Hong, Jingtong Dou, Chuancheng Shi, Yuhan Liu, Linrui Xu, Yixiong Liang, Yigang Cen, Yanpeng Sun, Fei Shen, Tat-Seng Chua
- 19Clinically-Grounded Hierarchical Classification for Consistent Chest X-ray InterpretationarXiv:2608.03016v1 Announce Type: new Abstract: Accurate chest X-ray interpretation is inherently hierarchical. Clinical decisions depend not only on what abnormality is present but where it is situated, requiring reasoning from broad anatomical systems down to specific pathological findings. Yet existing automated systems largely treat this as a flat classification problem, failing to capture inter-level dependencies or enforce coherence between coarse and fine predictions. We propose CHASE (ClJong Hak Moon, Minjun Kim, Minjun Kim
- 20Standalone DINOv3 for Training-Free Open-Vocabulary Semantic Segmentation in Remote SensingarXiv:2608.03023v1 Announce Type: new Abstract: Remote sensing semantic segmentation is hindered by costly pixel-level annotations, motivating training-free open-vocabulary methods. Recently, the recent release of DINOv3 brings DINO.txt, which equips the standalone DINO backbone with image-text contrastive learning and thus opens up the possibility of open-vocabulary segmentation. We propose DinoSplat-OV, a training-free framework that adapts DINOv3 to remote sensing without fine-tuning or additChanghao Zhao, Haoxiang Li, Yuke Li, Hai Liu, LingLin Zeng
- 21CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video GenerationarXiv:2608.03046v1 Announce Type: new Abstract: Text-to-video (T2V) diffusion transformers (DiTs) are trained with detailed video captions, whereas inference often relies on user prompts rewritten by a prompt enhancer (PE). Prior work has improved generation by optimizing the PE, the DiT, or both; some methods have also sought to narrow the training-inference mismatch through shared schemas. Yet even within a shared schema, inference-time PE outputs and DiT training captions may still differ inYizhuo Jia, Jingyun Hua, Yuanxing Zhang
- 22AIDE: Automated Instruction via Distilled Expertise for Reference-Free Motor Skill CoachingarXiv:2608.03047v1 Announce Type: new Abstract: Generating natural-language coaching feedback on motor skills can accelerate learning, yet expert coaches are scarce and expensive. Existing reference-based methods require expert demonstrations at both training and inference time, limiting practical deployment. We propose AIDE (Automated Instruction via Distilled Expertise), a framework that exploits expert references only during training and generates feedback from a learner's pose sequence aloneYoshiki Ito
- 23PDD-RRG: Posterior Diagnostic Decision for Study-level Radiology Report GenerationarXiv:2608.03055v1 Announce Type: new Abstract: Automatic radiology report generation (RRG) aims to simulate the workflow of radiologists, assisting them in clinical diagnosis. However, existing methods often fall short in utilizing all information relevant to the examination, as is typically done in clinical practice. Although some works attempt to incorporate multi-view images and historical data, these additional inputs may sometimes lead to avoidable diagnostic errors on the contrary. To addYang Yu, Yiming Ji, Bin Dai, Dong Zhang, Zhiyong Zhou, Shoushan Li, Yakang Dai
- 24TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion ModelsarXiv:2608.03057v1 Announce Type: new Abstract: Static quantization assigns one weight precision to every denoising step. To preserve quality, that precision must accommodate the most quantization-sensitive step, even though many other steps can tolerate fewer bits. The resulting model may satisfy its memory budget, but it repeatedly pays worst-case arithmetic throughout the denoising trajectory. We introduce Temporal-Adaptive Bit Sparsification Quantization (TASQ) to separate these two costs. TSeokho Han, Dongwei Wang, Jinhee Kim, Yiran Chen, Kang Eun Jeon, Huanrui Yang, Jong Hwan Ko
- 25RIDGE: Re-Noising with Internal Dynamic Guidance for Image EditingarXiv:2608.03059v1 Announce Type: new Abstract: Inversion-free flow-based image editing avoids latent inversion, but still requires a target-side state at every editing step. The widely used equal-displacement construction keeps the displacement between the noisy source state and the target-side state unchanged across noise levels. This is inconsistent with noising, under which the displacement between two clean states noised with the same noise level and noise sample should contract as the noisRuiliang Gong, Zhen Wang, Yanghao Wang, Long Chen
- 26Global Graph-Validated Optimization for VLM-based 3D Indoor Scene GenerationarXiv:2608.03064v1 Announce Type: new Abstract: We study open-vocabulary 3D indoor layout generation, which synthesizes diverse and physically plausible scenes from unlabeled 3D assets using free-form language instructions. Recent methods leverage large language models (LLMs) and vision-language models (VLMs) to generate structured scenes from text. However, most model inter-asset relations implicitly or rely on local pairwise constraints and local optimization. These formulations are poorly aliJialu Huang, Yingxuan You, Fei Wang, Zheng Dang
- 27LDU-Bench: Multimodal LLM Evaluation for Lithography Defect Understanding under Layout-Varying Circuit BackgroundsarXiv:2608.03078v1 Announce Type: new Abstract: Multimodal large language models have demonstrated strong defect recognition capability in industrial anomaly detection. However, in lithography review, merely determining whether an image contains a defect is insufficient for engineering inspection; models must also understand defect morphology, spatial location, and the potential causes supported by visible evidence. To this end, this paper proposes LDU-Bench, a multi-task multimodal benchmark foHuanglong Ji, Botong Zhao, Shujing Lv, Yue Lv
- 28CorePath: A Breast-Specialized Pathology Foundation Model for Core Needle Biopsy Diagnosis and Risk-Controlled Report GenerationarXiv:2608.03079v1 Announce Type: new Abstract: Breast core needle biopsy (CNB) is central to breast cancer diagnosis yet remains challenging because limited tissue sampling, lesion heterogeneity, and subtle morphologic overlap can obscure subtype distinctions. We developed CorePath, a breast-specialized multimodal pathology foundation model fine-tuned from PRISM using 7901 paired CNB whole-slide images and diagnostic reports from two centers. Evaluated across six CNB cohorts and two public breaTing Yin, Danning Li, Chen Shu, Xiaoxia Yao, Boyu Fu, Yujing Chang, Tianyu Shi, Mengna Feng, Jie Chen, Jing Fu, Xiuli Xiao, Tianlin Li, Mumin Shao, Jiaxin Bi, Wenchuan Zhang, Xiaoyan Wu, Xiao Han, Zhang Zhang, Yuhao Yi, Hong Bu
- 29DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion TransformersarXiv:2608.03082v1 Announce Type: new Abstract: Recent advances in Diffusion Transformers (DiTs) have enabled remarkable progress in visual synthesis, benefiting from their superior scalability. To facilitate DiTs' capability of capturing meaningful internal representations, recent works such as REPA incorporate external pretrained encoders for representation alignment. However, the underlying mechanisms governing representation learning within DiTs remain poorly understood in the community. ToBinglei Li, Mengping Yang, Zhiyu Tan, Xiaomeng Yang, Zhizhong Huang, Junping Zhang, Hao Li
- 30GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language ModelsarXiv:2608.03083v1 Announce Type: new Abstract: Video large language models (VideoLLMs) achieve strong video understanding performance, but their inference remains expensive due to the large number of redundant spatio-temporal visual tokens in long videos. Existing token pruning methods alleviate this cost by reducing redundant tokens, yet most of them rely on segment-level local pruning, where videos are partitioned into isolated segments and tokens are selected independently within each segmenMengjie Zhang, Qihui Zhu, Tao Zhang, Shuangwu Chen, Huihuang Qin, Yu Guo, Shenghao Ye, Zijian Wen, Yunpeng Hou, Dong Jin, Xiaobin Tan, Huasen He, Jian Yang



































