arXiv

知识

8榜单18天前更新默认榜单
cs.AI
arXiv
18天前更新
  • 01
    ISEE: Interactive Semantic Enrichment for Database Fields
    arXiv:2608.02604v1 Announce Type: new Abstract: LLM-based agents are increasingly being deployed for data-related tasks, including data sense-making, exploration, and retrieval. However, their performance heavily depends on the clarity and completeness of data semantics. In practice, many field descriptions remain ambiguous or incomplete, as much of the essential context (e.g., the meaning of a customized field) originates from users' domain knowledge and is rarely documented publicly. This gapYuan Tian, Yiru Chen, Rakesh R. Menon, Zifan Liu, Ting Cai, Fei Wu, Anudeep Chimakurthi, Prashanthi Ramamurthy, Sridevi Aishwariya Ganesan, Kun Qian, Yunyao Li
  • 02
    Self-Organising Digital Circuits
    arXiv:2608.02606v1 Announce Type: new Abstract: Fault tolerance in classical computing has traditionally relied on static strategies like hardware redundancy and error-correcting codes. Biological systems, in contrast, exhibit adaptive plasticity, maintaining function through dynamic re-organisation around damage. Inspired by this principle, we introduce Self-Organising Digital Circuits, framing functional logic generation and maintenance as a meta-learning problem on graphs. Our architecture emMarcello Barylli, Gabriel B\'ena, Alexander Mordvintsev, Eleni Nisioti, Sebastian Risi
  • 03
    Beyond the Hivemind: Escaping LLM Homogeneity via Meta-Persona Anchoring and Sequential Temperature Scaling
    arXiv:2608.02618v1 Announce Type: new Abstract: Recent studies have identified an ``Artificial Hivemind'' effect in Large Language Models (LLMs) causing models to converge on a narrow, homogenized consensus even for open questions. This semantic collapse limits the diversity of AI, resulting in high inter-response similarity ($\approx 0.80-0.90$) even under high-temperature sampling. In this paper, we propose a novel mitigation framework to increase diversity: Meta-Persona Anchoring combined witTairan Fu, Javier Conde, Carlos Arriaga, Gonzalo Mart\'inez, Pedro Reviriego, Javier Coronado-Bl\'azquez
  • 04
    PULSE: An Executable Contract Language for Spatiotemporal Knowledge Graph Engineering
    arXiv:2608.02630v1 Announce Type: new Abstract: Knowledge graph engineering often distributes accepted state, observations, constraints, processes, and hypothetical scenarios across artifacts whose combined execution contract remains external. We present PULSE, an Object-Process-Methodology-inspired language that localizes four operational roles and their write effects in one typed runtime. Here, modes denote operational roles rather than modal or deontic logic. The implemented contract fixes evDongxu Yang, Ziyi Liang
  • 05
    HyperAgent: Planning and Acting over Tool-Schema Hypergraphs for Tool-Use LLM Agents
    arXiv:2608.02650v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on external tools to complete complex real-world tasks. However, reliable tool-use planning remains challenging due to the limitations of implicit reasoning and the evolving nature of real-world execution environments. Existing tool-use agents typically rely on LLMs to infer tool compositions from textual descriptions, which can lead to inefficient exploration and unreliable execution in complex tZian Zhai, Xingyu Tan, Gaowang Zou, Xiaoyang Wang, Wenjie Zhang
  • 06
    Explainable AI for the EU Right to Explanation: A Systematic Review of the Law-XAI Translation Gap
    arXiv:2608.02699v1 Announce Type: new Abstract: When algorithms make or influence consequential decisions---about loan eligibility, hiring, or healthcare---EU law grants affected individuals a Right to Explanation. Yet whether (and how) Explainable AI (XAI) can satisfy this right in practice remains poorly understood, with direct implications for individuals' ability to contest automated decisions that affect their lives. This paper presents a systematic literature review of XAI in the context oBenjamin Fresz, Elena Dubovitskaya, Marco F. Huber
  • 07
    Predictive Set Theory: A Generative Framework for Cognitive Architecture with Operationalized Core Mechanisms
    arXiv:2608.02704v1 Announce Type: new Abstract: Predictive processing theories portray the brain as a hierarchical prediction engine that minimizes prediction error, yet they lack operational definitions for the structure of a "prediction," the standardized response to a prediction error, and the mechanism that maintains consistency across successive updates. Bayesian cognitive science attempts to subsume all uncertainty under probabilistic belief updating, but it presupposes a closed hypothesisYiyang Yu
  • 08
    Towards a new paradigm of scientific discovery with socialized artificial intelligence
    arXiv:2608.02775v1 Announce Type: new Abstract: Scientific discovery has advanced through successive transformations in the organization of knowledge. Observation and experimentation established the empirical foundations of science. Theory made it possible to derive general principles from particular phenomena. Computation extended inquiry into systems beyond direct observation, while data-intensive methods opened new spaces of pattern and prediction. Science now confronts a different frontier.Xinjie Yao, Xingxin Xu, Xiyuan Gao, Zhoupeng Guo, Kunlong Yang, Dengyu Zhao, Siqi Zhao, Zhihe Fan, Yichen Dong, Xin Li, Jiekang Feng, Jiahe Wu, Sen Wang, Beiming Yu, Kejia Zhao, Ruipu Zhao, Jiaqi Zhou, Heyang Li, Jianjun Chen, Anbo Dai, Xin Liu, Zhengtao Yu, Qinghua Hu, Pengfei Zhu
  • 09
    BAP-SQL: Budget-Aware Observation Planning for Agentic Text-to-SQL
    arXiv:2608.02876v1 Announce Type: new Abstract: Tool-using agents do not merely consume observations: their actions determine what arrives next. In agentic text-to-SQL, a broad query can spend context and database work before useful evidence appears, while post-hoc compression cannot recover omitted rows or expended work. We present BAP-SQL, which treats observation formation as a budget-control stage: it estimates query risk, rewrites SQL when useful, and delegates hard limits to an independentChong Peng, Pin Qian, Su Wang, Yihang Chen, Varun Sah
  • 10
    VeriTrace: Human-Like Temporal Exploration Completes Agentic Action Space
    arXiv:2608.02878v1 Announce Type: new Abstract: Large language models have shown promise for automated Verilog RTL generation, yet state-of-the-art multi-agent systems plateau at ~95% accuracy on standard benchmarks. We trace this ceiling to an incomplete debugging action space: existing systems restrict which signals the agent can inspect, which time windows it can query, or both, reducing debugging to pattern matching on a narrow, predetermined view of circuit behavior rather than hypothesis-dYu-Tung Liu, Cunxi Yu
  • 11
    Interpreting Black-Box Large Language Models with Sentence-Level Energy Landscapes
    arXiv:2608.02879v1 Announce Type: new Abstract: The widespread adoption of proprietary Large Language Models (LLMs) accessed strictly through closed APIs has created a critical challenge for responsible deployment: a fundamental lack of interpretability. To address this, we propose a model-agnostic, post-hoc attribution interpreter operating at the sentence level. Our approach trains an Energy-Based Model (EBM) as a surrogate to capture the LLM's internal conceptual consistency between prompts aMaryam Rezaee, Pooriya Safaei, Maryam Asgarinezhad, Fatemeh Seyyedsalehi
  • 12
    Hypercubes, Hyperplanes, and Constraint-Induced Complexity Collapse in Atomic Concept Learning
    arXiv:2608.02930v1 Announce Type: new Abstract: We revisit higher-arity atomic concept learning through the geometry of hypercubes and hyperplanes of ground instances. Our starting point is the observation that the ambient r-dimensional hypercube of ground atoms is not structurally uniform. Its logical complexity is organized by hyperplanes: every hyperplane other than the full diagonal collapses into finitely many elementary-equivalence classes, with a bound independent of the term depth, whileIrene Tsapara
  • 13
    When Compression Scores Cannot Decide: Information Boundaries for Group-Robust LLM Pruning
    arXiv:2608.02940v1 Announce Type: new Abstract: A reproducible compression statistic can still select the wrong candidate. A dense pruning score with 0.906 split-half reliability predicted a 16.1% gain. Its selected endpoint was 6.0% and 7.7% worse than two controls. We model the gap through information interfaces that delimit which distinctions each statistic supports. For equal-weight groups, a conic law gives the exact pooling price for positive linear fixed-candidate damage, including diagonAndrew Zhang
  • 14
    On the missing data layer and a potential solution
    arXiv:2608.02949v1 Announce Type: new Abstract: Latin America is missing two foundational layers of AI infrastructure: the dataset layer and the benchmark layer. This paper targets the dataset layer. The dataset layer faces two compounding problems: discovery and supply. Latin American AI datasets exist but are scattered across platforms with no shared index. Even with perfect indexing, the total volume would remain far below what frontier AI development requires. We propose DataHub: a task-firsFrancis F Daniel, Mauro Iba\~nez, Francis Perelman, Marian Basti
  • 15
    Neurosymbolic Reasoning with Incremental Knowledge for Sample Efficient Hierarchical Reinforcement Learning
    arXiv:2608.02993v1 Announce Type: new Abstract: (Flat) Reinforcement Learning (RL) agents face significant challenges in environments with sparse rewards that require long-horizon reasoning. A compelling approach to improve sample efficiency is to incorporate knowledge into learning and decision-making. In standard Hierarchical RL (HRL), knowledge is encoded in a fixed, non-updatable form, such as architectural choices, and remains unchanged throughout learning. With fixed HRL, reasoning with inSubrat Prasad Panda, Blaise Genest, Arvind Easwaran
  • 16
    On the missing benchmarks layer and a potential solution
    arXiv:2608.02996v1 Announce Type: new Abstract: Latin America is missing a foundational layer for native AI development: the benchmark layer. The benchmark layer does two things no other layer can - it audits AI systems against regional social requirements and it directs AI optimization in economically relevant environments. Without it, public institutions cannot independently evaluate foreign AI systems, and companies cannot optimize AI systems to solve local problems with SOTA performance. TheFrancis F Daniel, Mauro Iba\~nez, Francis Perelman, Marian Basti
  • 17
    ProPRL: Property-Aware Prerequisite Relation Learning in Educational Knowledge Graphs
    arXiv:2608.03006v1 Announce Type: new Abstract: Prerequisite relation learning is central to adaptive instruction, yet existing methods often formulate it as conventional link prediction, limiting their ability to adaptively integrate complementary educational evidence for individual candidate pairs and to discourage contradictory reverse predictions. We propose ProPRL, a Property-aware Prerequisite Relation Learning framework. ProPRL first learns complementary concept representations from a conXinghe Cheng, Jiapu Wang, Chaobo He, Ruihai Dong, Quanlong Guan
  • 18
    UrbanAgent: A Tool-Augmented Agent for Cross-System Urban Tasks
    arXiv:2608.03018v1 Announce Type: new Abstract: Modern cities rely on an increasing number of digital services to operate, but residents' daily needs are still difficult to meet. Services are fragmented and have little interoperability, placing a heavy operational burden on users. Existing digital platforms, urban foundation models, and intelligent assistants each address only isolated aspects of an urban task. But they struggle to reliably convert complex natural-language requests into executabJiayu Cao, Xingyuan Zeng, feiyu Li, Zhijing Huang, Xujie Yuan, Rongxiang Chen, Shimin Di, Libin Zheng, Jian Yin
  • 19
    LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment
    arXiv:2608.03020v1 Announce Type: new Abstract: Parameter-efficient post-training reduces the number of trainable parameters, but still requires repeated end-to-end backpropagation through the frozen backbone. Every adaptation step therefore needs backward-capable hardware and must store or recompute activations. We ask whether this repeated backward chain can be replaced by a one-time calibration. We introduce Local Credit Assignment (LoCA), a two-stage method for small-shift adaptation. One prLinhan Xia, Rui Liu, Zhaofeng Zhang, Yihao Wang, Binrui Shen, Shengxin Zhu
  • 20
    DiffImaginE: Imagine to Verify Entity Types with Diffusio
    arXiv:2608.03025v1 Announce Type: new Abstract: Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence. Existing imagine-and-compare verifiers map each (span, type) pair to one predicted visual feature, compressing diverse visual realisations into a single prototype and providing a compatibility score without explicit probabilistic semantics. We introduce DiffImaginE, which formulates MNER typFeng Zhang, Feiyu Han, Rongxin Yang, Yang Liu, Yancheng Chen, Rui Wang, Yingguang Yang, Tian Xueyun, Chongyang Zhang, Hao Zheng, Xu Kefu, Congjing Ran, Fuhai Chen, Bin Chong
  • 21
    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning
    arXiv:2608.03028v1 Announce Type: new Abstract: Applying a valid medication-safety rule when its patient-specific conditions are not met can produce an incorrect decision. Existing medical evaluations largely use isolated and fixed scenarios. A model may therefore answer correctly by recalling a drug-risk association without showing that it used patient information to decide whether the rule applies. To address this gap, we introduce MedPIC-Bench, a benchmark of source-verifiable recommendationsZhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang
  • 22
    CastFSR: A Fast--Slow--Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting
    arXiv:2608.03031v1 Announce Type: new Abstract: Time series forecasting is fundamental to decision-making in complex systems, where future dynamics are influenced not only by historical observations but also by evolving contextual features. Recent advances in large language models (LLMs) have extended forecasting beyond numerical extrapolation toward context-aware reasoning. However, existing approaches often lack explicit mechanisms to identify relevant contexts, reason about their impacts, andXiaoyu Tao, Mingyue Cheng, Bokai Pan, Chuang Jiang, Huanjian Zhang, Tian Gao, Yaguo Liu, Qi Liu, Enhong Chen
  • 23
    TraceCAD: Trace-Guided Repair for Agentic CAD Generation
    arXiv:2608.03062v1 Announce Type: new Abstract: LLM-based CAD agents produce executable parametric programs, but their correction loops may lose evidence about satisfied requirements, faulty operations, and prior repairs. We introduce TraceCAD, a recovery layer that links requested features, modeling steps, failure evidence, and candidate outcomes as persistent state. TraceCAD diagnoses likely faulty operations, searches bounded edits in their dependency regions, validates candidates through exeFengxiao Fan, Jingzhe Ni, Fan Sang, Xiaolong Yin, Yu Liu, Ruofeng Tong, Min Tang, Peng Du
  • 24
    Getting the Parameters Right: A Difficulty-Graded Benchmark and Probe-Guided Training for LLM Tool Calls
    arXiv:2608.03071v1 Announce Type: new Abstract: Large language model agents derive much of their capability from tool use. Existing research on tool use has largely focused on selecting the right tool and orchestrating the order of calls. However, correctly filling the parameters of a tool call is equally critical for successful execution and has received far less attention. In domains such as cloud networking, even frontier models correctly complete fewer than half of tool calls. Inspired by reGuoyao Yu, Xiaoqing Sun, Ziqi Huang, Shaojing Fan, Zhongyi Zhang, Xiaomeng Hu, Xiaobo Xue, Yangyang Shi, Xiong Xiao, Yang Song, Biao Lyu, Rong Wen, Xing Li, Qinming He, Shunming Zhu, Zhenguang Liu
  • 25
    AI Agent Economics: Can Autonomous Economic Behavior Emerge among AI Agents under Minimal External Conditions?
    arXiv:2608.03076v1 Announce Type: new Abstract: Multi-agent studies commonly place AI agents in predefined games, markets, or roles, making it difficult to distinguish endogenous economic organization from behavior inherited from the scenario. We ask whether economic relations emerge when agents receive executable mechanisms for work, transfer, elections, and allocation but no prescribed social or economic strategy. We define AI Agent Economics as systems of production, allocation, consumption,Lingyun Zhang, Shang Shang
  • 26
    Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR
    arXiv:2608.03119v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) improves LLM reasoning but typically relies on ground-truth (GT) answers, limiting scalability. Voting-based label-free RLVR replace gold supervision with answer-level consensus from model samples. However, collapse arises when the same answer-level signal is used both to estimate rewards and to drive token-level policy optimization, encouraging the model to directly reinforce answer tokens rathYongshi Ye, Liang Zhang, Yidong Chen, Xiaodong Shi, Biao Fu
  • 27
    Beyond Average Performance: Dynamic Instance Clustering and Specialized Algorithm Design in LLM-Assisted Evolutionary Search
    arXiv:2608.03129v1 Announce Type: new Abstract: Large Language Model-assisted Evolutionary Search (LES) has emerged as a powerful paradigm for automated algorithm design. However, existing LES methods primarily optimize for average performance, inherently directing search effort toward instances that contribute most to this metric while leaving others poorly served, resulting in weak tail robustness and limited real-world reliability. To address this limitation, we propose Dynamic Instance ClustQinglong Hu, Qingfu Zhang, Fei Liu, Xialiang Tong, Kun Mao, Mingxuan Yuan
  • 28
    Verifiable Memory: Learning Unified Memory Management with Local and Global Verifiers for Large Language Model Agents
    arXiv:2608.03137v1 Announce Type: new Abstract: Large language model (LLM) agents must retain reusable information, control a bounded active context, and recover earlier evidence during long-horizon interaction. Existing methods commonly optimize long-term memory (LTM) and short-term memory (STM) separately, while unified policies are often trained primarily with trajectory-level feedback, which provides weak credit for individual memory decisions. We present Verifiable Memory (VerMem), a framewXiaolong Sun, Qichao Wang, Hangyu Li, Liang Chen
  • 29
    Spatial proteomics guided by H&E-based AI reveals recurrence-risk niches in triple-negative breast cancer
    arXiv:2608.03145v1 Announce Type: new Abstract: Deep learning models can predict cancer recurrence from H&E stained slides, but the localized molecular states underlying these predictions remain largely obscured. Here, we developed an outcome informed spatial pathology framework in TNBC that integrates AI generated recurrence risk heatmaps with mass spectrometry based spatial proteomics. In a cohort of 156 patients, distribution based aggregation of high scoring patches achieved an AUC of 0.77 aYesung Cho, Ji Hwan Park, Chanil Kim, Hyewon Kim, Honglan Li, Yumin Lee, Geongyu Lee, Sujeong Hong, Seong Min Park, Yoonyoung Lee, Hee Sool Rho, Sumin Lee, Amos Chungwon Lee, Changhwan Lee, Hwanyoung Shim, Hyunwook Kim, Hyeji Shin, Sanha Park, Jihoon Yu, Yoon Hee Shin, Sooheon Kim, Hyunjin Park, Seung Min Park, Sangwan Kim, Yujung Kim, Sung-Im Do, Eun-Young Kim, Dongmyung Shin, Jongbae Park, In-Gu Do
  • 30
    UniGD: A Unified Generative-Discriminative Framework for Industrial Retrieval
    arXiv:2608.03150v1 Announce Type: new Abstract: Generative retrieval (GR) is a promising paradigm for industrial search advertising, yet its deployment is constrained by strict relevance and latency requirements. Existing systems cascade GR with an independent relevance model, decoupling the generative likelihood objective from query-ad relevance discrimination, which compromises effectiveness and increases serving costs. We propose a Unified Generative-Discriminative framework (UniGD) that inteShujie Ji, Yawei Kong, Yilin Zhao, Li Wang, Xialong Liu, Peng Jiang
cs.LG
arXiv
18天前更新
  • 01
    Deep Divide-and-Reduce in Symbolic Regression
    arXiv:2608.02628v1 Announce Type: new Abstract: Symbolic regression (SR) is the task of discovering underlying patterns from data and representing them using mathematical expressions. Current machine learning approaches to SR often lack a profound understanding of the intrinsic mathematical and physical principles governing these expressions. While the pioneering AI Feynman method leverages the mathematical properties underlying the data, its expression simplification mechanism suffers from a naYusong Deng, Yanjie Li, Weijun Li
  • 02
    Multimodal Auto-regressive Transformer Surrogate for Modeling Variable Operations and Quantifying Uncertainty in Geological Carbon Storage
    arXiv:2608.02629v1 Announce Type: new Abstract: The use of variable well perforation and injection strategies can improve the efficiency of geological carbon storage operations. We develop a new multimodal auto-regressive transformer surrogate to model these operations under geological uncertainty. A modified SEAM CO2 geomodel, which involves a faulted system with three stacked aquifers, is considered. The two injection wells are perforated in stages, from bottom to top, with the stage durationsYifu Han, Louis J. Durlofsky
  • 03
    LLMs Can Annotate Attribution Graphs
    arXiv:2608.02632v1 Announce Type: new Abstract: Circuit tracing is an exciting technique for revealing the internal computation of language models, but it requires a time-intensive manual step of grouping individual features or MLP neurons into supernodes. We present a simple pipeline for automating this step: directly presenting feature descriptions to a language model that groups them into supernodes. Using automated interpretability metrics, we confirm that supernodes generated by our pipelinAmeen Patel, Max Zhang, Nathan Hu
  • 04
    GeoID-PINN: Identifiability-Aware Regional Epidemic Inference with Geographic Coupling
    arXiv:2608.02633v1 Announce Type: new Abstract: Regional surveillance data reflect local transmission, reporting, seeding, and external infection pressure, which are difficult to identify separately. We introduce GeoID-PINN, a physics-informed neural network (PINN) for susceptible-infectious-recovered-deceased (SIRD) dynamics. The model represents spatial dependence with a row-stochastic source-composition matrix whose rows assign nonnegative source weights that sum to one. We regularize this maWeixiong Hua, Fan Bu
  • 05
    Verifier-Guided Model Discovery for Physical Dynamical Systems with Pretrained Symbolic Transformers
    arXiv:2608.02662v1 Announce Type: new Abstract: Reliable forecasting of nonlinear physical systems underpins scientific discovery and engineering decision-making. Yet high-fidelity simulations are prohibitively costly, and machine-learning surrogates can be opaque and encode assumptions about system dynamics, limiting generalizability. Pretrained transformers mapping synthetic ODE trajectories to equations offer interpretable alternatives, promising transfer without system-specific equation knowFarbod Faraji, Francesco Belardinelli
  • 06
    CT-HEG: A Bidirectional, Timestamp-Attributed Event Graph for ICU In-Hospital Mortality Prediction - An Architectural Ablation Study
    arXiv:2608.02663v1 Announce Type: new Abstract: Accurate ICU mortality prediction requires modeling irregular clinical observations across heterogeneous entity types. Existing sequence models handle irregular sampling but ignore typed relational structure; existing graph models assume fixed-interval inputs. We introduce the Continuous-Time Heterogeneous EHR Graph (CT-HEG) schema and evaluate which architectural choices drive predictive performance. CT-HEG encodes each ICU stay as a typed, timestMohammad Nasir Uddin, Rahnuma Tabassum Orpita, Asaduzzaman Anik, Eklachur Rahman Bhuiyan, Marjahan Risalat, SM Wali Ullah, Asif Ahamed
  • 07
    Sphere Retraction Normalizations
    arXiv:2608.02668v1 Announce Type: new Abstract: Residual connections are the de facto mechanism for training deep neural networks stably. Geodesic Normalization (GeoNorm) recasts them on a Riemannian manifold, orthogonalizing each layer output against the current hidden state and applying the resulting update through the Riemannian exponential map. Every hidden state thus keeps a constant $\ell_{2}$-norm, confining the residual stream to a hypersphere. The exponential map, however, is only one mJie Zhang, Cheng-Fang Su, Yi-Jui Huang, Min-Te Sun
  • 08
    Learning Molecular Representations from Cellular Phenotypes with Structure Preservation
    arXiv:2608.02688v1 Announce Type: new Abstract: Phenotypic drug discovery enables the discovery of functional relationships between molecular structures and cellular responses. However, existing multimodal representation learning methods often optimize cross-modal alignment without considering the intrinsic organization of chemical space, resulting in distorted molecular representations and loss of structural information. We propose \textbf{PhenMol}, a structure-preserving framework for phenotypXuan Lin, Jingyu Sheng, Tengfei Ma, Li Sun, Dapeng Xiong
  • 09
    GLOBE: Trajectory-Aligned Gradient Matching with Structured SparseOptimization for Coreset Selection
    arXiv:2608.02690v1 Announce Type: new Abstract: On-device training of deep neural networks is fundamentally constrained by the computational and memory costs of large-scale datasets. Coreset selection offers a practical solution by retaining only a compact subset of real training samples. However, existing gradient-based methods commonly rely on gradients computed at a single model snapshot and employ greedy or pursuit-based selection procedures, limiting their ability to capture evolving optimiHetian Liu, Jin Cui, Mengcheng Shi, Yanbin Hu, Xinyue Long, Boran Zhao, Pengju Pen
  • 10
    Output-Aware Rotation for INT2 KV-Cache Quantization
    arXiv:2608.02691v1 Announce Type: new Abstract: The key-value (KV) cache has become a major memory and bandwidth bottleneck in long-context large language model inference, making ultra-low-bit quantization increasingly important. However, existing rotation-based INT2 methods optimize cache statistics or proxy errors before the complete attention readout, even though the model is ultimately affected by the error propagated through attention and the output projection $W_O$. To address this mismatcVincent-Daniel Yun, Woosang Lim, Minsoo Cheong, Sunwoo Lee, Murali Annavaram, Sai Praneeth Karimireddy, Sungjoo Yoo
  • 11
    PatTree: a novel approach for automated creation of multimodal, graph-based patient representations for medical classification tasks
    arXiv:2608.02692v1 Announce Type: new Abstract: Access to holistic, multimodal data improves the performance of Artificial Intelligence (AI) in medical classification tasks compared to utilizing single modalities or data sources. However, the inherent heterogeneity and complexity of clinical real-world data pose significant challenges to structured data analysis and AI application. This heterogeneity includes missing values, multiple time points, diverse modalities, and inconsistent formats andJulia Gehrmann, Lars Quakulinski, Hamza Naseem, Oya Beyan
  • 12
    Measuring Explainer Stability via Attribution Separability
    arXiv:2608.02697v1 Announce Type: new Abstract: Attribution methods (AMs) assign an importance score to each feature and are widely adopted to explain black-box models. However, most methods can produce variable attribution scores due to stochastic components in their definition. In this paper, we propose a distribution-based framework to capture the stability of attribution scores. In particular, our approach allows to understand the degree of separability in the ranked attribution vector and oEddie Conti, \'Alvaro Parafita, Axel Brando
  • 13
    NANQ: Noise-Floor-Aware Mixed-Precision Non-Uniform Quantization for Analog Compute-in-Memory
    arXiv:2608.02700v1 Announce Type: new Abstract: Analog compute-in-memory (CIM) enables energy-efficient neural network inference, but device variation and read noise can severely degrade low-bit quantized models. Existing CIM-oriented quantization methods mainly minimize ideal quantization error, ignoring the hardware noise floor and thus causing inefficient precision allocation. We propose NANQ, a noise-aware mixed-precision non-uniform quantization framework for analog CIM. NANQ models magnituYizhe Chen, Wenshuai Yao, Saiya Wang, Yuannuo Feng, Wenbo Qi, Kechao Tang, Ngai Wong, Wenyong Zhou, Wang Kang
  • 14
    Can Training Logs Make Model Comparisons More Precise?
    arXiv:2608.02705v1 Announce Type: new Abstract: Comparing stochastically trained models requires estimating both a performance difference and its uncertainty from repeated runs. We study whether training logs from those same runs can make such comparisons more precise. Because training-log covariates are produced during training rather than measured before it, we use arm-specific covariate adjustment: each model is adjusted only with statistics from its own runs, and the raw mean difference remaWei-Jung Huang
  • 15
    Designing a Good Virtual Node: Addressable and Cardinality-Preserving Global Memory for Message Passing Architectures
    arXiv:2608.02709v1 Announce Type: new Abstract: Virtual nodes give message-passing neural networks a simple global communication route, but the standard node--VN--node pipeline compresses the graph into one homogeneous state and broadcasts it identically to every node. Building on the Two-Radius analysis of Mishayev et al., we ask how auxiliary virtual memory can relieve this finite-capacity bottleneck without self-attention. We identify two requirements. First, the global memory should be factoF\'elix Marcoccia
  • 16
    Neural Networks with Local Converging Inputs for Efficient Options Pricing Models
    arXiv:2608.02778v1 Announce Type: new Abstract: We present a novel application of Neural Networks with Local Converging Inputs (NNLCI) to improve the efficiency of existing numerical methods for pricing multi-asset options. The most concise input format for NNLCI has been introduced, offering substantial convenience and efficiency. NNLCI uses a neural network to locally correct solutions from a coarse mesh and a refined mesh (relative to the coarse one), requiring only a minimal amount of high-fHarris Cobb, Wenbo Hao, Yingjie Liu
  • 17
    Evaluation Blindness: How Silent Measurement Failures Corrupt AI Systems from Training to Deployment
    arXiv:2608.02786v1 Announce Type: new Abstract: AI systems can fail silently. The failure propagates through training loops, evaluation pipelines, and production monitoring stacks until downstream harm makes it visible. This paper introduces evaluation blindness: a measurement function M exhibits evaluation blindness with respect to failure class F when it produces readings indistinguishable from a healthy state while the system is actually failing, with no auxiliary signal flagging the gap. ThePriyanka Bajaj (Independent Researcher)
  • 18
    Topological Simplification in Predictive Coding Networks
    arXiv:2608.02816v1 Announce Type: new Abstract: We study the topology of learned representations in predictive coding networks (PCNs), a neuro-inspired bidirectional architecture, using a quantitative layer-wise persistent homology analysis. We train well-performing PCNs on a synthetic classification dataset ($\geq 99.9\%$ test accuracy) and on MNIST ($\geq 95\%$ test accuracy), and measure how topological features change across layers for different architectures and activation functions. We finAdam Shaw, Jiayu Li, Michael Sperling, Michael Kim, Alvin Jin
  • 19
    Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't
    arXiv:2608.02829v1 Announce Type: new Abstract: Model families train every size from scratch. Can a pretrained large model be converted into a smaller sibling? We characterize the 1.4B->410M conversion in the Pythia family end-to-end: (i) representations align strongly across sizes (ridge R^2=0.84) while parameters align weakly; (ii) dense weight projection is functionally destructive -- provably not an assembly artifact -- because basis mixing breaks rotary, per-head, GELU, and LayerNorm structRavi Satya Durga Prasad Yenugula
  • 20
    NOMADD: Numerical Optimization of Models Adapting to Data Drift
    arXiv:2608.02845v1 Announce Type: new Abstract: Tabular model performance degrades when feature distributions change over time or the relationship between features and outcome variables change over time, known as data drift and concept drift, respectively. These issues are challenging to mitigate in real time because labeled data may not be immediately available, or re-training a model could be impractical. While tools exist to reduce drift, they are typically bespoke to neural network architectSwapn Shah, Keith Burghardt
  • 21
    Adaptive Sampling for Automated Post-Disaster Rapid Damage Assessment via Level-Set Cost-Aware Bayesian Optimization
    arXiv:2608.02868v1 Announce Type: new Abstract: Natural disasters frequently inflict severe damage to the built environment, which demands a rapid, reliable, and cost-effective damage assessment for emergency response. However, traditional methods for post-disaster damage assessment often rely on static, labor-intensive data collection strategies that can be prohibitively expensive and struggle to adapt to dynamic post-disaster conditions. In this study, we propose a cost-aware Bayesian optimizaBoyang Xu, Mostafa Reisi Gahrooei, Mohammad Ilbeigi, Hao Yan
  • 22
    Contrast-invariant deep ptychography neural networks
    arXiv:2608.02869v1 Announce Type: new Abstract: Ptychography neural networks suffer from scaling inconsistencies when generalizing out of distribution, limiting their real world viability. We address this scaling mismatch using a factorization strategy which decouples the learned object texture from measurement scaling, enabling a single trained network to produce measurement-consistent reconstructions across varying illumination conditions. This requires predicting the learned object in real anAlbert Vong, Steven Henke, Oliver Hoidn, Hanna Ruth, Junjing Deng, Apurva Mehta, David Shapiro, Alexander Hexemer, Nicholas Schwarz
  • 23
    Maglev: Sliding Recurrent Memory
    arXiv:2608.02870v1 Announce Type: new Abstract: We introduce \ours{}, a recurrent Transformer architecture with fixed-size memory that generalizes sliding-window attention while remaining parallelizable during training. \ours{} consists of two coupled models: a prefiller $Q$, which leverages full attention\footnote{In practice, we use interleaved full and sliding-window attention for $Q$, as this yields stronger performance. The essential requirement is that $Q$ be more expressive than $P$, withBo Liu, Qiang Liu
  • 24
    GoT-CD: Graph-of-Thoughts Causal Discovery and the Fragility of Post-hoc Path-Specific Fairness Audits
    arXiv:2608.02877v1 Announce Type: new Abstract: Causal discovery recovers directed structure from observational data and is increasingly used in clinical settings to support mechanism reasoning and fairness audits of predictive models. Path-specific counterfactual fairness asks whether a protected attribute influences an outcome through illegitimate pathways, but these estimands are defined relative to a supplied causal graph and therefore inherit whatever errors the discovery step introduces. DNitish Nagesh, Elahe Khatibi, Thomas Dean Hughes, Mahdi Bagheri, Pratik Gajane, Amir M. Rahmani
  • 25
    Population-Robust Feature Selection via Generalized Welfare Optimization
    arXiv:2608.02887v1 Announce Type: new Abstract: Choosing which features to collect is a deployment decision: the same limited questionnaire, test panel, or sensor set may need to serve several heterogeneous populations. Standard feature-selection methods typically optimize for one large population, while existing robust approaches tend to learn one shared model for every population. We introduce PopFS, a method for learning one shared, deployable feature set that is robust to population differenRuiqi Lyu, Alistair Turcan, Bryan Wilder
  • 26
    Robust Counterfactual Policy Optimisation via Nondeterministic Causal Models
    arXiv:2608.02893v1 Announce Type: new Abstract: Counterfactual inference approaches for sequential decision-making typically assume deterministic causal models, where all randomness stems from latent variables. However, Markov Decision Processes (MDPs) are inherently stochastic. We address this by formalising counterfactual policy optimisation under probabilistic nondeterministic causal models, which properly separates latent confounding from irreducible stochasticity, and here propose a first pJessica Lally, Milad Kazemi, Nicola Paoletti, David Watson, Sander Beckers
  • 27
    AnchorKV: Anchor-Residual KV Cache Compression
    arXiv:2608.02901v1 Announce Type: new Abstract: The key-value (KV) cache is the primary memory bottleneck in long-context LLM inference. Existing approaches attack it from opposite ends: eviction methods permanently discard tokens, degrading performance whenever a discarded token later proves essential, while quantization methods retain all tokens at low precision but offer limited compression. We propose AnchorKV, a compression scheme that shrinks the cache by $20\times$ without discarding a siMalik Khalaf, Yara Shamshoum, Nitzan Hodos, Yuval Sieradzki, Assaf Schuster
  • 28
    Bayesian Data Reweighting Improves Multimodal Retrieval for Knowledge-Based Visual Question Answering
    arXiv:2608.02907v1 Announce Type: new Abstract: Multimodal retrievers are essential for knowledge-based visual question answering, where they retrieve external evidence for image-question pairs. However, existing contrastive training methods typically treat all unmatched query-document pairs as equally informative negatives, which is problematic because many unmatched documents may still be semantically relevant or partially useful. We propose Bayesian Data Reweighting, a probabilistic frameworkJingchen Sun, Shaobo Han, Ruiyi Zhang, Naresh Kumar Devulapally, Ming Liu, Yitao Long, Vishnu Suresh Lokhande, Changyou Chen
  • 29
    Forecasting Revenue with its Customer-Base Drivers: When and Why Coordination Helps
    arXiv:2608.02911v1 Announce Type: new Abstract: Revenue forecasts guide acquisition budgets, demand planning, and customer-based valuations, yet an aggregate forecast does not show whether change reflects acquisition, repeat purchasing, spending per order, or offsetting movements. Using weekly transaction panels for 966 companies in 25 industries, the authors develop the Customer-Based Multi-task Transformer (CBMT), which learns shared structure, retains separate primitive forecasts, and alignsKyeongbin Kim, Daniel McCarthy, Dokyun Lee
  • 30
    When Should Graph Attention Be Sparse? Learning a Per-Edge Tsallis Index
    arXiv:2608.02938v1 Announce Type: new Abstract: Graph attention normalizes neighborhood scores with softmax, the maximum-entropy choice under Shannon statistics. But homophilic and heterophilic graphs want different attention shapes, and one fixed normalization cannot serve both. We propose \textbf{LTGA} (\textbf{L}earnable \textbf{T}sallis \textbf{G}raph \textbf{A}ttention), a graph attention layer whose Tsallis entropic index $q$ is learned jointly with the weights, interpolating continuouslyKleyton da Costa, Bernardo Modenesi
cs.CL
arXiv
18天前更新
  • 01
    TabletCraft: Bridging a 4,000-Year Cultural Gap with Bidirectional Akkadian NMT and Cuneiform Rendering
    arXiv:2608.02609v1 Announce Type: new Abstract: Half a million cuneiform clay tablets survive in museums worldwide, yet modern users can neither read nor write in the world's oldest writing system, leaving a 4,000-year cultural barrier that existing NLP tools have only partially addressed. Prior work enables one-way, scholar-oriented translation from Akkadian to English, but offers no path in the reverse direction: non-specialist users cannot compose new content in cuneiform, and therefore remaiZhaohui Wang
  • 02
    BBOWP-Bench: Evaluating LLMs on Black-Box Optimization Word Problems
    arXiv:2608.02612v1 Announce Type: new Abstract: Formulating an optimization problem strongly affects the quality of the final solution, yet good formulations usually require substantial expertise. Recent studies have therefore examined how to automatically derive optimization problems from natural-language descriptions, but existing benchmarks focus on settings where objectives and constraints can be written explicitly as mathematical expressions. Many practically important problems are naturallYutaro Yamada, Kei Hiroshima, Nozomu Yoshinari, Kento Uchida, Shinichi Shirakawa
  • 03
    MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale
    arXiv:2608.02613v1 Announce Type: new Abstract: Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models. Yet, existing memory benchmarks often under-test the combination of activity-dense interaction, ego-centric perspective, and coherent multi-session worlds. MemArena fills these gaps with a single-world conversational benchmark built with its MASim agent simulator, for 50 agents over 15 days (10.3M dialog-text tokens, 24.1K texJiadong Zhang, Xiaosong Ma
  • 04
    OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning
    arXiv:2608.02615v1 Announce Type: new Abstract: Cancer diagnosis and characterization require integrating complementary evidence from radiology, pathology, genomics, and clinical metadata. However, most medical large language model (LLM) and vision-language model (VLM) benchmarks focus on isolated modalities or narrow image-text tasks, leaving patient-level oncology assessment across multiple evidence streams largely untested. We introduce OncoTriad-QA, a patient-level radiology-pathology-genomiAhnaf Munir, Dannong Wang, Michael W. McDonald, Mubarak Shah, Pegah Khosravi, Yu Tian
  • 05
    Evaluating OpenAI's Privacy Filter: Cross-Lingual, Cross-Domain PII Detection Across 42 Benchmarks
    arXiv:2608.02616v1 Announce Type: new Abstract: We present the first independent, systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-parameter bidirectional PII detector, across 42 synthetic benchmarks spanning 22 languages and 5 domains. Zero-shot, OPF achieves F1=0.855 on AI4Privacy and 0.464 on SPY medical, outperforming Presidio (0.431, 0.273) and XLM-RoBERTa (0.269, 0.111) on PII-annotated benchmarks; on multilingual NER, XLM-RoBERTa leads OPF on all 13 Indic and non-Latin languRohith Uppala
  • 06
    Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety
    arXiv:2608.02617v1 Announce Type: new Abstract: We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive Open Online Validation and Evaluation), a clinician-led platform collecting blinded pairwise preferences alongside multi-criterion rubric ratings. Clinicians assign scores on a discrete $[-2, +2]$ scale, where negative values indicate clinically unsafe or misleading contenFay Elhassan, David Sasu, Alexandra Kulinkina, Lars Henning Klein, Mary-Anne Hartley
  • 07
    JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation
    arXiv:2608.02620v1 Announce Type: new Abstract: LLM-as-a-judge evaluation has become a dominant paradigm for ranking language models, yet the ecosystem remains fragmented: most benchmarks ship their own code base, hardcode a specific closed-model judge, and support a single evaluation protocol. This fragmentation makes it difficult to study how design choices--the benchmark, the judge model, the prompt, the inference backend--affect the conclusions we draw about model quality. We introduce JudgeErlis Lushtaku, Bora Kargi, Ali Elganzory, Fabio Ferreira, Alejandro R. Salamanca, Julia Kreutzer, David Salinas
  • 08
    Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks
    arXiv:2608.02621v1 Announce Type: new Abstract: Legal benchmarks typically score final answers even when models also state legal authority. We test whether answer correctness can serve as a proxy for authority grounding. Under ordinary reasoning prompts that did not request statutory citations, four LLMs spontaneously produced authority markers across 238 Taiwan bar-examination items. Because each item has a verified governing provision, we automatically audit answer correctness and authority grHsien-Jyh Liao
  • 09
    Speculative Correction: Draft-then-Refine Decoding for Diffusion Language Models
    arXiv:2608.02625v1 Announce Type: new Abstract: Diffusion language models (DLMs) can revise tokens bidirectionally, but standard decoding procedures often adapt them to left-to-right generation by producing text block by block. We study a simple plug-and-play inference pattern: first generate a complete draft, then refine the full response using bidirectional diffusion. Using LLaDA2.1-Flash and LLaDA2.1-Mini, we evaluate two configurations. In Flash-Flash, the same Flash model serves as both draBrian K Chen, Chong Wu, Kenji Kawaguchi
  • 10
    Stuck on "A": Diagnosing and Repairing Interface Injury in Attention-to-KDA Linearization of a 0.6B Language Model
    arXiv:2608.02689v1 Announce Type: new Abstract: We convert 21 of 28 full-attention layers of Qwen3-0.6B-Base into KDA (Kimi Delta Attention) linear-attention layers on a single consumer-grade GPU budget, and ask a simple question: what exactly does the conversion break? After surgery, hidden-state alignment and end-to-end KL distillation drive the student close to its teacher in perplexity, yet multiple-choice accuracy stays near random chance (25-29% vs. the teacher's 50.6% on C-Eval). Using aRonglong Bao
  • 11
    Crayotter: Learning Long-Horizon Video Editing Agents via Group-Relative Preference Backpropagation
    arXiv:2608.02694v1 Announce Type: new Abstract: Long-horizon video editing agents receive final-product feedback only after many interdependent decisions. Yet editing quality is subjective, admits multiple valid solutions, and is not meaningfully calibrated across heterogeneous requests, making a global scalar objective both ambiguous and temporally uninformative. Our key observation is that fixing the request, materials, and production constraints converts this subjective objective into an ordiLecheng Yan, Jianze Lin, Yichong Zhang, Ben Pan, Wenxi Li, Chenyang Lyu, Liting Zhou, Cathal Gurrin
  • 12
    ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads
    arXiv:2608.02703v1 Announce Type: new Abstract: Weight-only quantization substantially reduces the storage of large language model (LLM) transformer blocks, but practical backends often retain the final language-modeling head (LM-head) in BF16 or FP16. Quantizing this projection naively can strongly perturb the vocabulary-logit distribution. We present ARCHead, a packed LM-head compressor that combines a quantized low-rank core, group-wise INT4 residuals, and a low-rank correction fitted in an a\c{S}uayp Talha Kocabay, Talha R\"uzgar Akku\c{s}, Kamer Ali Yuksel
  • 13
    Learning a Vector-Symbolic Model for Socio-Cultural Tasks
    arXiv:2608.02807v1 Announce Type: new Abstract: How can we better represent the impact of sociocultural structures on decision making in computational cognitive models? Modeling this impact requires traversing multiple levels of semantic representation, however it is not immediately clear to a modeler which levels of representation are most salient to a given situation. Though large language models and cognitively grounded corpus models can represent broad semantic associations through co-occureMeera Ray, Swapnika Dulam, Christopher L. Dancy
  • 14
    BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?
    arXiv:2608.02867v1 Announce Type: new Abstract: Although reinforcement learning with verifiable rewards (RLVR) has improved the performance of large language models (LLMs) across a variety of reasoning tasks, there is significant debate as to whether RLVR expands the reasoning capability boundary, or just improves sampling efficiency. In this paper, we investigate the nature of test-time exploration in RLVR-trained LLMs by employing controlled maze-solving experiments and extracting a tree strucSoumadeep Saha, Krish Sharma, Akshay Chaturvedi, Nicholas Asher
  • 15
    FLARE: Few-shot Learning-based Adaptive Reflective Engine
    arXiv:2608.02919v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in complex, compound AI systems where performance hinges on the quality of prompts. Recent state-of-the-art optimizers like GEPA (Genetic-Pareto) have argued that reflective instruction evolution can outperform traditional reinforcement learning and few-shot optimization. In this work, we challenge this shift by introducing FLARE (Few-shot Learning-based Adaptive Reflective Engine), a frameworkDhanasekar Sundararaman, Bharat Gandhi, Aashna Garg, Minjie Li
  • 16
    Character Iconicity vs. Arbitrariness: An Arabic NLP Perspective
    arXiv:2608.02935v1 Announce Type: new Abstract: Arabic script uses 28 letters, many of which share a common base shape (rasm) and are distinguished only by dot placement. Because early Arabic manuscripts were written without dots yet remained interpretable, dot removal offers a natural test of whether these visual distinctions are functionally necessary. Prior work has shown that dotless Arabic can remain readable and effective for natural language processing (NLP), but it remains unclear whetheDorieh Alomari, Irfan Ahmad, Maged S. Al-shaibani
  • 17
    Aligned in Form, Not in Meaning: The Comprehension - Containment Decoupling of LLM Safety in Low-Resource Bangla Derogatory Speech
    arXiv:2608.02941v1 Announce Type: new Abstract: We audit five frontier large language models on native Bangla derogatory speech (gali) across six protocols to test a single hypothesis: Comprehension-Containment Decoupling. We propose that contemporary safety alignment is bound to high-resource surface forms rather than harmful meaning, causing a model's capacity to comprehend a low-resource slur and its capacity to contain it to operate independently. Every protocol corroborates this hypothesisShadab Bin Habib, A K M Ferdous Reza Habib, Subarno Neel, Adib Sakhawat
  • 18
    OPTD: On-Policy Transition Distillation with Consistency-Guided Adaptive Compression for Few-Step Diffusion Language Models
    arXiv:2608.02942v1 Announce Type: new Abstract: Diffusion language models (dLLMs) can predict many tokens in parallel, but accurate generation still requires many iterative denoising steps. Few-step distillation accelerates decoding by compressing multiple teacher steps into a single student transition. However, existing methods construct supervision on off-policy trajectories. At inference, the student's early parallel commitments alter the context of later predictions, so the states it actuallXiaocheng Lu, Hualei Zhang, Shuhan Guo, Jie Zhang, Xiaoyi Pang, Jian Liu, Haoxi Li, Bohai Gu, Haoxuan Che, Jingcai Guo, Song Guo
  • 19
    Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks
    arXiv:2608.02966v1 Announce Type: new Abstract: Most multiple-choice question (MCQ) benchmarks evaluate Large Language Models (LLMs) only by whether they select the correct answers. This binary scoring treats all incorrect responses alike, even though an LLM's preferences among incorrect options may contain systematic and useful information about its behavior and ability. We introduce the LLM Nominal Response Model (LLM-NRM), an option-aware psychometric framework that models the full distributiXiao Fei, Yang Zhang, Sarah Almeida Carneiro, Michalis Vazirgiannis
  • 20
    Mapping the City Through the Lens of Language Models
    arXiv:2608.02971v1 Announce Type: new Abstract: Language models often complete an underspecified reference to a city with unstated assumptions about urban size, form, infrastructure, environment, and function. We measure those assumptions without naming places. Ten open-weight checkpoints rate anonymized profiles derived from real morphological urban centres across 40 audited indicators and seven domains. The design combines constrained probability-based ratings, prespecified reliability screensWanqi Liu, Rong Zhao, Zhizhou Sha, Qinyu Cui, Yecheng Zhang
  • 21
    TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation
    arXiv:2608.02975v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated impressive performance in MQM-based translation quality (TQ) evaluation, and recent advances in large reasoning models (LRMs) promise even greater improvements. However, both LLMs and LRMs are computationally expensive to deploy at scale, while small language models (SLMs)---though much more efficient---struggle with the complex reasoning required for evaluation tasks. In this work, we present an extenBhavin Jawade, Cameron R. Wolfe
  • 22
    On the Non-Specificity of Statistical Measures Used in Script Decipherment
    arXiv:2608.02999v1 Announce Type: new Abstract: Statistical regularities are routinely offered as evidence that undeciphered sign systems encode language; the Indus script debate is the canonical example. Any such inference rests on specificity: the reported outcome must be unusual among plausible structured non-languages. We test that premise constructively with SIGIL, a purpose-built generative emblem system whose 3,000-text core corpus carries explicit compositional meanings although no signNikhil Raghavendra
  • 23
    Language Models Encode the Contextual Truth of Propositions
    arXiv:2608.03035v1 Announce Type: new Abstract: Prior work has shown that LLMs encode the truth of factual propositions along linear directions in activation space. It's unclear how these representations extend to contextual truth: propositions whose truth is determined by in-context evidence rather than world knowledge. We show that LLMs maintain a linear representation of contextual truth that persists across structurally different output policies, even when the output doesn't require the modeRupak Sarkar, Pritika Ramu, Rachel Rudinger
  • 24
    Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models
    arXiv:2608.03038v1 Announce Type: new Abstract: Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how models construct and communicate statistical explanations. This study demonstrates the value of a multidimensional evaluation by combining response accuracy, response behavior, structural topic modeling, and lexical similarity analysis. The framework is applied to explanations generated by 15 currentMonnie McGee, Mateo Langston Smith, Julian Cabrera
  • 25
    Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation
    arXiv:2608.03044v1 Announce Type: new Abstract: Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignment with human survey data, while others find persona collapse and weak demographic sensitivity. We show that much of this conflict stems from conflating two distinct tasks. We call the first task emulation, in which models generate individual responses that aggregate into a population distribution. WSeth Grief-Albert, Jessica Bo, Difan Jiao, Ashton Anderson
  • 26
    PI-Mem: Pushing Long-Context Reasoning to 3.6M Tokens with Parallel-Iterative Memory
    arXiv:2608.03048v1 Announce Type: new Abstract: Long-context reasoning remains a critical bottleneck for large language models, as recent recurrent-memory approaches face two inherent challenges: sequential chunk-wise updates can overwrite early critical evidence with later irrelevant content, and serial inter-chunk dependencies limit parallelism and cause latency to increase with context length. To address these issues, we propose PI-Mem (Parallel-Iterative Memory), a mechanism that processes aDawei Liu, Haixu Song, Shuang Cheng, Shijie Wang, Haozheng Hou, Kaifeng Liu, Ermo Hua, Zhonghang Yuan, Zhijie Zhong, Yuchen Fan, Biqing Qi, Bowen Zhou
  • 27
    SeqLLM: Augmenting LLMs with Behavioral-Sequence Modeling for High-Stakes Decisions at WeChat Pay
    arXiv:2608.03063v1 Announce Type: new Abstract: Merchant risk control at large payment platforms screens tens of millions of merchants daily, where false positives harm legitimate merchants and false negatives leave harmful activity undetected. The hardest cases require jointly understanding a merchant's textual profile and long behavioral sequence. Large language models (LLMs) excel at text but cannot natively model such sequences, while adapting them often causes catastrophic forgetting. We prGuilin Li, Jiaxing Zhang, Matthias Hwai Yong Tan, Bo Wang, Weiran Huang
  • 28
    Activation-Guided Neuron Intervention to Induce Alzheimer's-Related Computational Language Phenotypes in a Large Language Model
    arXiv:2608.03067v1 Announce Type: new Abstract: Changes in spontaneous speech provide an early signal of cognitive dysfunction in Alzheimer's disease (AD) that large language models (LLMs) can detect. However, detection alone cannot establish whether the underlying model representations contribute functionally to behavior. We introduce an activation-guided intervention framework using Qwen3-8B. The framework identifies feed-forward neurons with higher activation rates for AD than control transcrRui He, Ercong Nie, Hong Jiang, Iris E. Sommer, Philipp Homan, Wolfram Hinzen
  • 29
    CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning
    arXiv:2608.03068v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as an effective method for enhancing the reasoning capabilities of large language models (LLMs). However, existing methods suffer from insufficient precision in feedback on generated answer trajectories and exhibit the phenomenon of problem difficulty drift. To address these challenges, we propose CVPO - Curriculum-guided Value-Variance Policy Optimization. At the response trajectory level, we find that tokenZiqi Jia, Yalu Ouyang, Bo Pang, Panpan Li, Hangfei Xu, Shengzhao Wen, Shiyong Li, Yanpeng Wang
  • 30
    PAMT: Process-Aligned Reinforcement Learning for Multi-Domain Machine Translation
    arXiv:2608.03077v1 Announce Type: new Abstract: Multi-domain machine translation (MDMT) requires more than fluent generation: it demands domain-sensitive translation decisions such as domain disambiguation, terminology control, and stylistic adaptation. Large reasoning models (LRMs) make such decisions explicit through intermediate translation steps, but our analysis across 15 domains and four translation directions shows that this explicit reasoning is double-edged: it improves long-form and hiYongshi Ye, Biao Fu, Chongxuan Huang, Yidong Chen, Xiaodong Shi
cs.RO
arXiv
18天前更新
  • 01
    Light-Loco-Parkour: Versatile Perceptive Whole-Body Locomotion via Multi-Skill Distillation
    arXiv:2608.02653v1 Announce Type: new Abstract: Existing humanoid whole-body control systems still fall short of the way humans move through cluttered terrain: they either track expressive whole-body references without terrain generalization, or react to terrain online while leaving the arms, torso, and knees largely unused. We present \texttt{Light-Loco-Parkour} (LLP), an end-to-end perceptive whole-body locomotion system that closes this gap with a single deployable policy. Conditioned only onHongming Chen, Zhuoran Li, Hongxi Wang, Jiangpeng Hu, Ziliang Li, Peize Liu, QingRui Zhao, Xuhao Liu, Liang Pan, Ximin Lyu, Yuntao Ma, Tingxiang Fan
  • 02
    Semantic Haptic Feedback Enhances Dexterous Robotic Teleoperation
    arXiv:2608.02780v1 Announce Type: new Abstract: In robot teleoperation, haptic feedback can be used to help human operators accomplish dexterous manipulation tasks. However, existing haptic feedback methods try to replicate high-fidelity sensory haptics that are felt in real world interactions, which are constrained by the sensing and feedback hardware capability and may lead to higher workload. To addresses these limitations, this work introduces semantic haptics for teleoperation, which uses aBingjian Huang, Sahar Aseeri, Jonas Schmidtler, Joseph Zhang, Sonny Chan, Andrew Doxon, Jom Preechayasomboon, Evan Pezent, Alberto Rigo, Amir Memar, Nicholas Colonnese, Chase Tymms
  • 03
    Toward Certified Functional Safety for Industrial Humanoid Robots: The Fail-Passive Gap and a Feasibility Study
    arXiv:2608.02809v1 Announce Type: new Abstract: Industrial humanoid robots are constrained less by locomotion or manipulation capability than by the immaturity of functional safety certification for legged platforms. The root difficulty is that the safe state of a legged robot is an actively-controlled state, which violates the fail-passive assumption underlying ISO~13849-1 / EN~60204-1: removing power from a walking biped causes an uncontrolled fall, so classical de-energization is itself a hazCaiwu Ding, Tao Cui, Lingyun Wang, Chengtao Wen
  • 04
    Staying on Spec: Real-Time Monitoring under Uncertainty with a Maritime Case Study
    arXiv:2608.02811v1 Announce Type: new Abstract: Robotic systems must operate under uncertainty while satisfying complex task and safety specifications. Monitoring such specifications under uncertainty remains challenging, as existing formulations typically require extensive data or explicit uncertainty distributions. In this paper, we propose a real-time monitoring framework that reduces data requirements by leveraging data-driven reachable sets for specification evaluation. We instantiate the fElizabeth Dietrich, Hanna Krasowski, Emir Cem Gezer, Roger Skjetne, Asgeir Johan S{\o}rensen, Murat Arcak
  • 05
    Biconvex Optimization for Smooth Minimum-Time Trajectories around Convex Obstacles
    arXiv:2608.02834v1 Announce Type: new Abstract: We present a biconvex approach for minimum-time motion planning around convex obstacles that is guaranteed to converge, is anytime, and supports derivative constraints to arbitrary order. We jointly convexify the minimum-time objective and all derivative constraints through a change of variables, and handle collision avoidance via time-varying separating planes, reducing the problem to a biconvex program. This program is solved by alternating betwePeter Werner, Tobia Marcucci, Daniela Rus
  • 06
    Control Barrier Functions via Minkowski Operations for Safe Navigation among Polytopes
    arXiv:2608.02886v1 Announce Type: new Abstract: Safely navigating polytopic environments while respecting the dynamics, control, and exact geometry of the underlying system is a challenge in robotics. Control barrier functions (CBFs) synthesize safe control policies by rendering the safe set forward invariant, but many existing CBF-based methods approximate polytopes using conservative smooth shapes, such as spheres or ellipsoids, to obtain explicit differentiable distance functions. In this artYi-Hsuan Chen, Shuo Liu, Wei Xiao, Calin Belta, Michael Otte
  • 07
    Contact-Driven Localization in a Freeform Robotic Self-Assembled Structure
    arXiv:2608.02895v1 Announce Type: new Abstract: Accurate localization remains a key challenge in swarm robotics, particularly for self-reconfigurable systems that must identify relative positions to form diverse structures. Most existing approaches rely on external tracking infrastructure or high-cost sensors, which limit scalability and deployment in unstructured environments. In this paper, we propose a novel contact-driven localization method for modular robots that leverages only local commuMohammadali Rashidioun, Michael Sosa, Petras Swissler
  • 08
    DeRP: An Algorithm for Self-Assembly of Power-Delivery Networks using Recursive Branching in Information-Limited Environments
    arXiv:2608.02904v1 Announce Type: new Abstract: Delivering sustained power to distributed equipment in unstructured field environments using pre-planned wired networks or battery-based solutions presents significant infrastructure and logistics challenges. This paper presents Dendritic Recursive Pivoting (DeRP), a decentralized framework for multi-target network formation in robot swarms based solely on local communication and bearing-based sensing toward sinks. We envision a system in which robMohammadali Rashidioun, Sangwoo Park, Petras Swissler
  • 09
    ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies
    arXiv:2608.02958v1 Announce Type: new Abstract: Vision-Language-Action (VLA) policies trained by behavior cloning fail silently: from the action stream alone, a collapsing rollout looks much like one making clean progress, because imitation supplies no notion of progress. Reinforcement learning would supply one, but it is impractical here, where real-robot experience is costly and deformable food resists simulation. The cheap alternative, a terminal success / failure bit, is learnable in principInkyu Sa, Konstantin Stulov, Rajat Bhageria
  • 10
    EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation
    arXiv:2608.02990v1 Announce Type: new Abstract: Latent diffusion models (LDMs) have recently significantly advanced embodied learning in constructing powerful embodied manipulation world models. However, despite the remarkable performance, existing LDMs predominantly rely on Variational Autoencoders (VAEs) optimized for natural scenes while failing to account for the unique characteristics of embodied manipulation scenarios, yielding latent representations that are neither compact nor controllabJiayi Luo, Hanxin Zhu, Chen Gao, Jiankun Wang, Cong Wang, Tianyu He, Jianxin Li, Zhibo Chen
  • 11
    A Wearable Stiffness-Rendering Haptic Device with a Honeycomb Jamming Mechanism for Bilateral Teleoperation
    arXiv:2608.03002v1 Announce Type: new Abstract: This paper addresses the challenge of providing kinesthetic feedback in bilateral teleoperation by designing a wearable, lightweight (20 g), and compact haptic device, the HJ-Haptic, utilizing a honeycomb jamming mechanism for object stiffness rendering. The HJ-Haptic device can vary its stiffness, from 1.15 N/mm to 2.64 N/mm, using a 30 kPa vacuum pressure. We demonstrate its implementation in a teleoperation framework, enabling operators to adjusThomas M. Kwok, Bohan Zhang, Wai Tuck Chow
  • 12
    Forbidden Region Dynamic Active Constraints in Robot-Assisted Minimally Invasive Surgery
    arXiv:2608.03010v1 Announce Type: new Abstract: In robot-assisted surgery, Forbidden Region Active Constraints (FRAC) represent a control strategy that helps maintain task safety by generating anisotropic haptic guidance to surgeons. However, several challenges need to be overcome before FRAC can benefit teleoperative surgery in a clinical setting. These challenges include the ability to allow for dynamic tissue deformation, maintain energetic passivity, and speed of implementation, among othersZejian Cui, Ferdinando Rodriguez y Baena
  • 13
    PACE: Adaptive Budget Allocation for Time-Efficient Embodied Planning
    arXiv:2608.03034v1 Announce Type: new Abstract: Reasoning-enhanced large language models have achieved remarkable improvements in planning tasks, yet their deployment in embodied systems remains impractical due to prohibitive inference delays-often exceeding minutes per planning instance. The fundamental bottleneck stems from the serial nature of existing paradigms: models must complete all reasoning before any action execution, leaving execution time windows entirely unexploited. We introduce PYuchen Huang, Xijiang Ying, Zhenhua Ma, Xiaxiang Yuan, Zhijie Gao, Jiayi Huang, Ruichi Mao, Jiazheng Zhang, Hongsheng Ti, Maotao Tian, Rong Shi, Lu Zhao, Shizhuang Zhang, Zhuo Cui, He Wang, Ling Liu, Wei Zhang
  • 14
    CUDA MPC: A GPU-Native Solver for Model Predictive Control
    arXiv:2608.03051v1 Announce Type: new Abstract: Model Predictive Control (MPC) delivers constraint-aware control, but its reliance on online optimization limits its use on systems with fast dynamics, high-dimensional models, or long horizons. Existing GPU implementations typically treat the device as a linear-algebra accelerator, leaving the optimization loop dependent on repeated kernel launches and high-latency memory transfers. This paper introduces CUDA MPC, a GPU-native MPC framework that cBabak Akbari, Melissa Greeff
  • 15
    How Should Vision-Language-Action Models Use Proprioceptive State?
    arXiv:2608.03052v1 Announce Type: new Abstract: Recent Vision-Language-Action (VLA) models almost universally take robot proprioceptive state as input, yet wire it in incompatible ways -- serialized into text prompts, projected into the vision-language prefix, or fed directly to the action expert -- and almost always as a single current frame. Three questions remain open: (1) whether, and on which tasks, current state actually improves closed-loop control; (2) how much state history helps, and wYiren Zhao, Ziyang Chen, Ziyang Rao, Pengteng Li, He Zhang, Weiyu Guo, Yandong Guo, Rushi Dai
  • 16
    Passively Safe Convex Guidance for Cislunar Rendezvous and Proximity Operations
    arXiv:2608.03060v1 Announce Type: new Abstract: This paper presents purely convex programs for passively safe impulsive rendezvous and proximity operations in cislunar orbits. Approach, arrival, and abort maneuvers are all designed and validated in the context of maneuver execution error and navigation uncertainty, and formulated for efficient onboard execution in the autonomous scenario. The outlined methods form the baseline onboard guidance routines for NASA's CAPSTONE 02 mission planned to dIan M. Down, Connor Plaks, Matthew Bolliger, Michael Caudill
  • 17
    A Hierarchical Approach to Imitation Learning for Manipulation Tasks Requiring Time Varying Forces
    arXiv:2608.03103v1 Announce Type: new Abstract: Diffusion policies have shown strong performance in learning complex, multi-modal behaviors for robotic manipulation. However, their application to contact-rich disassembly tasks remains limited by a key trade-off: the iterative denoising process introduces inference latencies that makes high frequency control difficult, which is essential for realizing dynamic interactions such as chiseling and prying. Recent action-chunking techniques mitigate laRishabh Shukla, Adithya Santhosh, Shaili Gandhi, Samrudh Moode, Satyandra K. Gupta
  • 18
    Shooting for Contact: Contact-Implicit Multiple Shooting for Dynamic Motion Retargeting
    arXiv:2608.03116v1 Announce Type: new Abstract: Motion retargeting approaches often prioritize kinematic similarity over whole-body dynamics, contact consistency, and actuation limits, yielding references that are difficult for reinforcement learning (RL) policies to reproduce, particularly for contact-rich behaviors. We present a contact-implicit, direct simulation-based multiple shooting (DSMS) framework that transforms kinematically feasible references into dynamically feasible whole-body traSergio A. Esteban, Jason H. K. Siu, Derrick Mach, Junheng Li, Vince Kurtz, Joel W. Burdick, Aaron D. Ames
  • 19
    DigitCode: Symbolic Tokenization of Hand Motion by Anatomical Units
    arXiv:2608.03127v1 Announce Type: new Abstract: Hand motion carries the finest-grained information in human activity, yet the representations behind hand generation, understanding, and robot learning are overwhelmingly continuous--joint angles or MANO parameters. These are accurate but unstructured: a finger cannot be indexed or edited as a symbol, and nothing marks a pose as anatomically valid. Discrete symbolic representations supply exactly this structure, and Hand Labanotation (HL) has shownHaoyu Gu, Haotian Lu, Jingrun Du, Xiao-Ping Zhang
  • 20
    POMDPs for Autonomous Science Exploration
    arXiv:2608.03155v1 Announce Type: new Abstract: Autonomous exploration missions require decision-making under sensor uncertainty and computational constraints, yet integrating scientific representations into POMDP planning has remained intractable due to high-dimensional observation spaces. Information-theoretic planners overcome this by assuming deterministic observations, sacrificing the principled uncertainty quantification that POMDPs provide. We introduce the Science Hypothesis Map POMDP (SDaniel Guirguis, Nathan Wallace, Hanna Kurniawati, Salah Sukkarieh
  • 21
    Accelerating Human-Aware Robot Trajectory Generation via Diffusion and Consistency Distillation
    arXiv:2608.03159v1 Announce Type: new Abstract: This research proposes a constrained motion planning framework for robot manipulators in human-robot interaction (HRI). For a non-redundant manipulator with a fully specified end-effector pose, additional requirements such as collision avoidance and self-collision avoidance are difficult to handle as simple null-space secondary tasks. This limitation makes it challenging to generate feasible joint-space trajectories in HRI environments where safetyByeong-Il Ham, Hyun-Bin Kim, Kyung-Soo Kim
  • 22
    PFM-HR: Pose Flow Matching for Humanoid Robots
    arXiv:2608.03227v1 Announce Type: new Abstract: Motion priors improve reinforcement learning for physics-based humanoid tracking, but temporal priors require ordered motion clips, while pose priors provide limited guidance for policy-induced pose transitions. We present Pose Flow Matching for Humanoid Robots (PFM-HR), a reusable flow matching prior trained directly on large scale unordered pose data. PFM-HR introduces the Pose Geometry Score (PGS), which quantifies how joint coordinate changes dYukang Gao, Yi Gu, Yangchen Zhou, Xingyu Chen, Zhaorui Wang, Fanghai Zhang, Hanyang Cao, Zhengyang Shen, Ji Ma, Runhan Zhang, Lei Han, Renjing Xu
  • 23
    Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking
    arXiv:2608.03231v1 Announce Type: new Abstract: Vision-Language-Action (VLA) policies promise general robotic manipulation, but their robustness against physical-world attacks remains fragile. In particular, we show that physically realizable adversarial patches can reliably induce failures by triggering a mechanism we call policy-critical action-to-vision attention hijacking, where action-conditioned attention is diverted from task-relevant regions to a localized patch. To demonstrate the threaJinquan Zhang, Dongfu Yin, Run Yang, Yufeng Yan, Zhen Tian, F. Richard Yu
  • 24
    Learning Context-Aware Motion Priors for Humanoid Control
    arXiv:2608.03234v1 Announce Type: new Abstract: Motion priors provide powerful guidance for learning naturalistic humanoid behaviors. However, existing methods typically learn a general, task-agnostic prior from the entire reference dataset and apply it uniformly throughout policy training. As a result, the prior cannot distinguish which reference motions are relevant to the current task context, potentially providing irrelevant or conflicting guidance. We present Context-Aware Motion Priors (CMYunyang Mo, Yi Gu, Yangchen Zhou, Hanyang Cao, Renjing Xu
  • 25
    GraspMeanFlow: SE(3)-Equivariant MeanFlow for Few-Step 6-DoF Grasp Generation
    arXiv:2608.03295v1 Announce Type: new Abstract: Recent data-driven methods for synthesizing 6-DoF grasp poses use generative models to learn complex grasp pose distributions and generate diverse candidate poses. In particular, SE(3)-equivariant flow-based models generate grasp poses that transform consistently with object rotations and translations. However, these methods sample by iterative numerical integration, requiring tens of function evaluations per grasp and limiting their use in real-tiJiyong Kwon, Yikun Bai, Amirhossein Mollaali, Guang Lin
  • 26
    PLS-Calib: A Partial Least Squares Framework for Event Camera and Odometry Calibration under Ground Motion Constraints
    arXiv:2608.03296v1 Announce Type: new Abstract: Accurate extrinsic rotation calibration between sensors is fundamental to the performance of robotic perception systems. However, most existing calibration techniques rely on full 6-DoF motion to excite all degrees of freedom, which is often infeasible for ground-constrained robots with limited motion capabilities. Recent approaches designed for such restricted settings, such as Canonical Correlation Analysis (CCA)-based methods, suffer from ill-coGuangyu Li, Xiao Li, Yujie Wu, Changshuo Wang, Prayag Tiwari, Jiang Cai, Fangwen Yu, Mingkun Xu
  • 27
    Shaping Wind-Tunnel Airflow for Unmanned Aerial Vehicles using Online Learning
    arXiv:2608.03378v1 Announce Type: new Abstract: The development and testing of advanced aerial robots require experiments in controlled environments with tailored airflow profiles. This paper presents an online learning algorithm for controlling the complex airflow field in a multi-fan vertical wind tunnel. Our method combines a simplified physical model with iterative, measurement-based learning, enabling sample-efficient convergence to desired airflow distributions. We demonstrate the method'sGhadeer Elmkaiel, Michael Muehlebach
  • 28
    RoboReact: Agentic Skill Distillation from Generated Egocentric Videos for Generalizable Whole-Body Manipulation
    arXiv:2608.03387v1 Announce Type: new Abstract: Humanoid robots have the potential to perform dexterous manipulation in human environments, yet acquiring diverse and generalizable skills remains costly due to expensive hardware data collection and labor-intensive annotation. Recent advances in video generative models provide a promising opportunity to synthesize rich manipulation experiences from visual observations, but transferring such imagined behaviors into executable whole-body humanoid skShuliang He, Shuai Wang, Bo Yue, Junchi Teng, Changyu Wang, Guiliang Liu
  • 29
    Flying over The Uncertain Nature (FORTUNE): Intelligent and Humanistic 3D Path Planning for Low-Altitude Collaboration
    arXiv:2608.03408v1 Announce Type: new Abstract: The proliferation of low-altitude intelligent agents is increasing the demand for timely and socially responsible collaborative sensing in dynamic urban environments. However, jointly addressing heterogeneous spatiotemporal demands, environmental uncertainty, and human-centered operational constraints remains challenging. This paper studies 3D multi-UAV path planning and task assignment under uncertain ground PoI demands. Unlike existing work assumMinghui Liwang, Wenhan Jia, Xinlei Yi, Wenbo Zhu, Yuhan Su, Xianbin Wang
  • 30
    A Low-Cost Hybrid Reservoir Computing Model for Isolated Sign Language Video Recognition
    arXiv:2608.03444v1 Announce Type: new Abstract: Sign language recognition (SLR) enhances communication between hearing and hearing-impaired individuals. Although deep learning (DL) has achieved promising performance in SLR, its high computational cost limits deployment on edge devices. To address this challenge, we propose a lightweight reservoir computing (RC)-based approach for SLR. In the proposed method, MediaPipe extracts body and hand keypoints to capture the spatial and temporal dynamicsNitin Kumar Singh, Arie Rachmad Syulistyo, Yuichiro Tanaka, Hakaru Tamukoh
cs.CV
arXiv
18天前更新
  • 01
    Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
    arXiv:2608.02711v1 Announce Type: new Abstract: Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimodal data, particularly the lack of large-scale and geometrically consistent editing data. To address this limitation, we propose Hunyuan3D-Buffalo 1.0, a unified framework supporting 3D understanding, text-to-3D generation, instructioJunliang Ye, Kenkun Liu, Guocun Wang, Yang Li, Yansong Qu, Chunshi Wang, Jingwei Xu, Yunhan Yang, Zibo Zhao, Jiachen Xu, Jiaao Yu, Lifu Wang, Zhihao Liang, Zhuo Chen, Chunchao Guo
  • 02
    Quo Vadis, World Modeling?
    arXiv:2608.02713v1 Announce Type: new Abstract: Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real-environment interaction is costly, slow, unsafe, and hard to parallelize. World modeling offers a natural intermediate proxy that allows agents to query lower-cost, more controllable feedback before committing to real actions. Classical world models instantiate this proxy primarily through future physical-state prediction, a formulation usefYu Yang, Xuemeng Yang, Licheng Wen, Lingdong Kong, Xiaobin Hu, Dongyue Lu, Wei Chow, Xiyan Huang, Yuxiang Feng, Yue Liao, Jianbiao Mei, Daocheng Fu, Rong Wu, Pinlong Cai, Ran Yi, Ying Tai, Jiangning Zhang, Botian Shi, Yong Liu, Shuicheng Yan
  • 03
    Oh Deer, How Should I Handle This? Seasonal Priors for Selective Wildlife Annotation and Classification
    arXiv:2608.02762v1 Announce Type: new Abstract: Fine-grained wildlife classification in aerial imagery is limited not only by model performance, but also by unreliable labels: animals occupy few pixels, key visual cues vary seasonally, and modality-specific evidence can be ambiguous. We study adult-male identification in red deer, where the antler cycle defines predictable windows of reliable evidence for both annotation and prediction. Using 7,295 RGB-only, thermal-only, and matched RGB+thermalHugo Markoff, Christoph Praschl, Anton Hjalte J{\o}rgensen, Christian Emil Mogensen, Mathias Bech Skadhauge, Sara Beery, Michael {\O}rsted, David C. Schedl
  • 04
    Confident but Unreliable: A Behavioral Safety Audit of Vision-Language Models on Brain MRI
    arXiv:2608.02790v1 Announce Type: new Abstract: Vision-language models (VLMs), including medical specialists, are increasingly proposed for medical imaging, yet their stated confidence is rarely evaluated separately from correctness. We use brain MRI as a controlled, high-stakes testbed for a broader failure mode in frontier multimodal systems: models can appear competent while lacking reliable self-knowledge. We present an automatically graded behavioral audit and pilot study of six instructionAmir Sabbaghziarani, Mohammadsajad Abavisani, Sergey Plis
  • 05
    Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation
    arXiv:2608.02791v1 Announce Type: new Abstract: MLLM-based segmentation faces a core segmentation trilemma: high segmentation performance, preserved dialogue ability, and fast inference. Embedding-prediction methods may disrupt language modeling through pixel-level objectives, whereas next-token generation is inefficient for dense masks. We propose All-Mask Prediction, decoupling autoregressive dialogue from non-autoregressive mask prediction. Its binary instantiation, STAMP (Simultaneous TextuaJiazhen Liu, Mingkuan Feng, Long Chen
  • 06
    PixelUp: Zero-Shot Semantic Feature Upsampling for Fine-Grained Vision Tasks
    arXiv:2608.02792v1 Announce Type: new Abstract: Self-supervised Vision Foundation Models (VFMs) have become essential backbones for downstream tasks due to their strong and transferable visual representations. However, their patch-token-level features are often too coarse for dense prediction tasks such as semantic segmentation and depth estimation when accurate fine-grained predictions are required. Feature upsampling methods have been developed to recover pixel-level detail but still face limiDeepank Singh, Anurag Nihal, Vedhus Hoskere
  • 07
    SAGE: Semantic Explainability of Attention-Based Survival Models in Computational Pathology
    arXiv:2608.02803v1 Announce Type: new Abstract: Attention-based multiple instance learning (ABMIL) is the predominant approach for slide-level prediction in computational pathology, yet its attention maps provide only local explanations: they indicate where a model focuses but not which histological features drive its predictions or how the model behaves across a patient cohort. We present Semantic Attention Global Explanations (SAGE), a post-hoc framework that extracts global, language-groundedAbdallah Lamane, Abdul Rahman Diab, Ren-Chin Wu, William Lotter
  • 08
    A Unified 2D Framework for DeepLesion Detection, Segmentation and Short Report Generation
    arXiv:2608.02805v1 Announce Type: new Abstract: In previous work, we integrated large language models (LLMs) into the lesion segmentation model based on the ULS23 DeepLesion dataset, using short-form findings from the reports. In this study, we developed a unified 2D lesion analysis framework that integrates LLM-based reasoning, lesion bounding box detection, segmentation, and radiology report generation from the original DeepLesion dataset. In the testing phase, we achieved relatively high lesiRuida Cheng, Tejas S. Mathai, Benjamin Hou, Qingqing Zhu, Zhiyong Lu, Matthew McAuliffe, Ronald M. Summers
  • 09
    In-Context Collapse in Vision-Language Models and How to Mitigate it?
    arXiv:2608.02830v1 Announce Type: new Abstract: Many-shot in-context learning (ICL) lets vision-language models (VLMs) adapt from image--label demonstrations without weight updates, and is widely assumed to improve as more demonstrations are supplied. We show the opposite: as demonstrations accumulate, a subset of VLMs undergo an \emph{in-context collapse}, a sharp, sometimes catastrophic accuracy drop spanning synthetic classification, natural-image classification, and VQA benchmarks, in some mMohammad Rostami
  • 10
    CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning
    arXiv:2608.02833v1 Announce Type: new Abstract: Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension with logical reasoning, yet current models struggle with accurate visual grounding and coherent reasoning chains. While extrinsic chain-of-thought prompting and visual cues significantly improve performance, current MLLMs lack intrinsic visual grounded reasoning capabilities, leading to inaccurate perception and reasoning disconnectedXuehang Guo, Pingyue Zhang, Ruiyi Zhang, Zhenhailong Wang, Hanrui Lyu, Heng Ji, Tong Sun, Qingyun Wang, Manling Li
  • 11
    A Human-in-the-Loop Deep Learning Framework for Color Reconstruction of Lenticular Films
    arXiv:2608.02835v1 Announce Type: new Abstract: Historical lenticular films, such as those created with the Kodacolor process, encode color information in a distinctive spatial format. This structure requires specialized techniques for accurate color reconstruction. While recent signal processing approaches like doLCE and deep learning methods like deep-doLCE have advanced automated color recovery, they often fail with cases such as curved lenticules, low-contrast, or badly captured regions. WeSaptarshi Neil Sinha, Tiago Kleist, Giorgio Trumpy
  • 12
    Localize, Don't Beautify: Client-Side Control of Image-Editing APIs for Cosmetic Surgery Previews
    arXiv:2608.02841v1 Announce Type: new Abstract: Ask a commercial image editor to preview a cosmetic procedure and it will often change more of the face than the request names: a nose edit can also smooth skin or alter lighting. Existing methods for confining an edit to one region require access to the model's internals, which a public editing API does not expose. We ask how much control is possible from the client side alone. In a pilot benchmark, six commercial editing configurations and one maSukhrobbek Ilyosbekov
  • 13
    Test Time Adaptation Methods for Point Cloud Registration in Laparoscopic Surgery
    arXiv:2608.02883v1 Announce Type: new Abstract: 3D point cloud registration in laparoscopic surgery estimates the transformation between an intraoperative organ reconstructed from video and its preoperative mesh. Because ground-truth transformations are unavailable for real data, supervised networks are trained on synthetic organ pairs. At test time, real reconstructions differ from synthetic data and are noisy, sparse, and occluded, which degrades correspondence estimation. Test-time adaptationNina Bodelot, Soufiane Belharbi, Eric Granger
  • 14
    Modeling Scientific Experiment Scenes: Dataset and Model
    arXiv:2608.02892v1 Announce Type: new Abstract: Scene Graph Generation (SGG) is fundamental to structured visual understanding, yet existing benchmarks focus mainly on daily life images and overlook scientific experiment scenes with specialized instruments, task-specific experimental semantics, and dense, fine-grained physical relations. These scenes are increasingly important for automated experimental analysis and smart education. To bridge this gap, we introduce PhysScene, the first SGG datasMinghao Zou, Qingtian Zeng, Shangkun Liu, Cong Liu, Paul L. Rosin, Guanghui Yue, Jun Liu, Wei Zhou
  • 15
    RealWeather: Realistic and Scene-Faithful Weather Translation with Driving World Models
    arXiv:2608.02953v1 Announce Type: new Abstract: Realistic weather translation is valuable for developing and evaluating autonomous driving systems, yet collecting paired videos of the same scenes under different weather conditions at scale is impractical. Existing methods therefore rely on synthetic data, 3D weather editing, or geometry-conditioned generation, often compromising weather realism or scene fidelity. We propose RealWeather, a driving world model for both realistic and scene-faithfulYuwei Ning, Liangzhi Wang, Yi Xiao, Zhenhua Wu, Yun Pang, Mingkun Chan, Jichang Li, Guanbin Li
  • 16
    Material-Segmented Per-Pixel Emissivity Correction for Thermographic Anomaly Detection in Cultural Heritage Digital Twins
    arXiv:2608.02964v1 Announce Type: new Abstract: Quantitative longwave thermography of heritage surfaces is limited by the global-constant emissivity assumption in inverse-Planck temperature retrieval; on heterogeneous surfaces emissivity varies within one field of view, producing apparent-temperature artifacts that mimic and mask subsurface anomalies. We present a training-free pipeline that derives per-pixel emissivity by applying SAM 3.1 open-vocabulary segmentation to a colocated, co-calibratJonathan Klingspon, Scott McAvoy, Maurizio Seracini, Falko Kuester
  • 17
    Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding
    arXiv:2608.02980v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) have achieved remarkable success on images and short videos, yet scaling them to long videos remains challenging due to frame-centric tokenization and limited context windows. 3D geometry provides a natural compression mechanism for visual streams: depth and camera pose enable observations from multiple views and time steps to be fused into a persistent, world-aligned representation. While recent 3D LMMs leverage geomLucy Lin, Ayush Jain, Yifan Liu, Katerina Fragkiadaki
  • 18
    V-FIND: Revealing the Intrinsic Forgery Knowledge Encoded in Video Forgery Detectors
    arXiv:2608.03008v1 Announce Type: new Abstract: As generated videos become increasingly realistic, reliable video forgery detection is increasingly important. Existing studies typically optimize and use video forgery detectors as black boxes, while the latent forgery-discriminative knowledge inside them remains largely unexplored. Instead of continuing to rely on resource-intensive full-model retraining to steadily improve detection performance, we ask whether video forgery detection can also beShichao Kan, Chengpeng Hong, Jingtong Dou, Chuancheng Shi, Yuhan Liu, Linrui Xu, Yixiong Liang, Yigang Cen, Yanpeng Sun, Fei Shen, Tat-Seng Chua
  • 19
    Clinically-Grounded Hierarchical Classification for Consistent Chest X-ray Interpretation
    arXiv:2608.03016v1 Announce Type: new Abstract: Accurate chest X-ray interpretation is inherently hierarchical. Clinical decisions depend not only on what abnormality is present but where it is situated, requiring reasoning from broad anatomical systems down to specific pathological findings. Yet existing automated systems largely treat this as a flat classification problem, failing to capture inter-level dependencies or enforce coherence between coarse and fine predictions. We propose CHASE (ClJong Hak Moon, Minjun Kim, Minjun Kim
  • 20
    Standalone DINOv3 for Training-Free Open-Vocabulary Semantic Segmentation in Remote Sensing
    arXiv:2608.03023v1 Announce Type: new Abstract: Remote sensing semantic segmentation is hindered by costly pixel-level annotations, motivating training-free open-vocabulary methods. Recently, the recent release of DINOv3 brings DINO.txt, which equips the standalone DINO backbone with image-text contrastive learning and thus opens up the possibility of open-vocabulary segmentation. We propose DinoSplat-OV, a training-free framework that adapts DINOv3 to remote sensing without fine-tuning or additChanghao Zhao, Haoxiang Li, Yuke Li, Hai Liu, LingLin Zeng
  • 21
    CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generation
    arXiv:2608.03046v1 Announce Type: new Abstract: Text-to-video (T2V) diffusion transformers (DiTs) are trained with detailed video captions, whereas inference often relies on user prompts rewritten by a prompt enhancer (PE). Prior work has improved generation by optimizing the PE, the DiT, or both; some methods have also sought to narrow the training-inference mismatch through shared schemas. Yet even within a shared schema, inference-time PE outputs and DiT training captions may still differ inYizhuo Jia, Jingyun Hua, Yuanxing Zhang
  • 22
    AIDE: Automated Instruction via Distilled Expertise for Reference-Free Motor Skill Coaching
    arXiv:2608.03047v1 Announce Type: new Abstract: Generating natural-language coaching feedback on motor skills can accelerate learning, yet expert coaches are scarce and expensive. Existing reference-based methods require expert demonstrations at both training and inference time, limiting practical deployment. We propose AIDE (Automated Instruction via Distilled Expertise), a framework that exploits expert references only during training and generates feedback from a learner's pose sequence aloneYoshiki Ito
  • 23
    PDD-RRG: Posterior Diagnostic Decision for Study-level Radiology Report Generation
    arXiv:2608.03055v1 Announce Type: new Abstract: Automatic radiology report generation (RRG) aims to simulate the workflow of radiologists, assisting them in clinical diagnosis. However, existing methods often fall short in utilizing all information relevant to the examination, as is typically done in clinical practice. Although some works attempt to incorporate multi-view images and historical data, these additional inputs may sometimes lead to avoidable diagnostic errors on the contrary. To addYang Yu, Yiming Ji, Bin Dai, Dong Zhang, Zhiyong Zhou, Shoushan Li, Yakang Dai
  • 24
    TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion Models
    arXiv:2608.03057v1 Announce Type: new Abstract: Static quantization assigns one weight precision to every denoising step. To preserve quality, that precision must accommodate the most quantization-sensitive step, even though many other steps can tolerate fewer bits. The resulting model may satisfy its memory budget, but it repeatedly pays worst-case arithmetic throughout the denoising trajectory. We introduce Temporal-Adaptive Bit Sparsification Quantization (TASQ) to separate these two costs. TSeokho Han, Dongwei Wang, Jinhee Kim, Yiran Chen, Kang Eun Jeon, Huanrui Yang, Jong Hwan Ko
  • 25
    RIDGE: Re-Noising with Internal Dynamic Guidance for Image Editing
    arXiv:2608.03059v1 Announce Type: new Abstract: Inversion-free flow-based image editing avoids latent inversion, but still requires a target-side state at every editing step. The widely used equal-displacement construction keeps the displacement between the noisy source state and the target-side state unchanged across noise levels. This is inconsistent with noising, under which the displacement between two clean states noised with the same noise level and noise sample should contract as the noisRuiliang Gong, Zhen Wang, Yanghao Wang, Long Chen
  • 26
    Global Graph-Validated Optimization for VLM-based 3D Indoor Scene Generation
    arXiv:2608.03064v1 Announce Type: new Abstract: We study open-vocabulary 3D indoor layout generation, which synthesizes diverse and physically plausible scenes from unlabeled 3D assets using free-form language instructions. Recent methods leverage large language models (LLMs) and vision-language models (VLMs) to generate structured scenes from text. However, most model inter-asset relations implicitly or rely on local pairwise constraints and local optimization. These formulations are poorly aliJialu Huang, Yingxuan You, Fei Wang, Zheng Dang
  • 27
    LDU-Bench: Multimodal LLM Evaluation for Lithography Defect Understanding under Layout-Varying Circuit Backgrounds
    arXiv:2608.03078v1 Announce Type: new Abstract: Multimodal large language models have demonstrated strong defect recognition capability in industrial anomaly detection. However, in lithography review, merely determining whether an image contains a defect is insufficient for engineering inspection; models must also understand defect morphology, spatial location, and the potential causes supported by visible evidence. To this end, this paper proposes LDU-Bench, a multi-task multimodal benchmark foHuanglong Ji, Botong Zhao, Shujing Lv, Yue Lv
  • 28
    CorePath: A Breast-Specialized Pathology Foundation Model for Core Needle Biopsy Diagnosis and Risk-Controlled Report Generation
    arXiv:2608.03079v1 Announce Type: new Abstract: Breast core needle biopsy (CNB) is central to breast cancer diagnosis yet remains challenging because limited tissue sampling, lesion heterogeneity, and subtle morphologic overlap can obscure subtype distinctions. We developed CorePath, a breast-specialized multimodal pathology foundation model fine-tuned from PRISM using 7901 paired CNB whole-slide images and diagnostic reports from two centers. Evaluated across six CNB cohorts and two public breaTing Yin, Danning Li, Chen Shu, Xiaoxia Yao, Boyu Fu, Yujing Chang, Tianyu Shi, Mengna Feng, Jie Chen, Jing Fu, Xiuli Xiao, Tianlin Li, Mumin Shao, Jiaxin Bi, Wenchuan Zhang, Xiaoyan Wu, Xiao Han, Zhang Zhang, Yuhao Yi, Hong Bu
  • 29
    DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformers
    arXiv:2608.03082v1 Announce Type: new Abstract: Recent advances in Diffusion Transformers (DiTs) have enabled remarkable progress in visual synthesis, benefiting from their superior scalability. To facilitate DiTs' capability of capturing meaningful internal representations, recent works such as REPA incorporate external pretrained encoders for representation alignment. However, the underlying mechanisms governing representation learning within DiTs remain poorly understood in the community. ToBinglei Li, Mengping Yang, Zhiyu Tan, Xiaomeng Yang, Zhizhong Huang, Junping Zhang, Hao Li
  • 30
    GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models
    arXiv:2608.03083v1 Announce Type: new Abstract: Video large language models (VideoLLMs) achieve strong video understanding performance, but their inference remains expensive due to the large number of redundant spatio-temporal visual tokens in long videos. Existing token pruning methods alleviate this cost by reducing redundant tokens, yet most of them rely on segment-level local pruning, where videos are partitioned into isolated segments and tokens are selected independently within each segmenMengjie Zhang, Qihui Zhu, Tao Zhang, Shuangwu Chen, Huihuang Qin, Yu Guo, Shenghao Ye, Zijian Wen, Yunpeng Hou, Dong Jin, Xiaobin Tan, Huasen He, Jian Yang
eess.SY
arXiv
18天前更新
  • 01
    Coordinated Primary Frequency Regulation and Grid-Forming Control for Wind Turbine Generators
    arXiv:2608.02813v1 Announce Type: new Abstract: Conventional grid-forming (GFM) control strategies often treat the DC source as an unconstrained link, creating mismatches when applied to the wind turbine generators (WTGs). Focusing on primary frequency regulation, this paper systematically investigates the mismatch between the GFM-WTGs behavior and the droop-based primary frequency regulation. To address this issue, a novel coordination strategy between WTG primary frequency regulation and GFM cMeng Chen, Yufei Xi, Lin Cheng, Florian D\"{o}fler, Ioannis Lestas
  • 02
    Stabilization of First-Order Partial Integro-Differential Equations with Concurrent Input and State Delays
    arXiv:2608.02851v1 Announce Type: new Abstract: This paper considers boundary stabilization problems for a first-order hyperbolic partial integro-differential equation (PIDE) subject to concurrent input and state delays. The coexistence of these two types of delays complicates control design, especially under the case of large input delay that requires to predict more state information. A backstepping-based boundary controller is developed to achieve stabilization and delay compensation. The desSanguan Zhong, Jie Qi
  • 03
    Sequential Operational Decision-Making for Power System Resilience Under Evolving Wildfires
    arXiv:2608.02976v1 Announce Type: new Abstract: This paper proposes a novel automated decision-support framework aimed at enhancing the resilience of power systems and operational resilience against wildfires by formulating the decision-making process as a stochastic multi-stage programming during a progressive wildfire. The paper develops a framework that takes into account both preventive and corrective actions, enabling automated and adaptive decisions based on potential scenarios over the coArastoo H Salimi, Majid Dehghani, Hamidreza Nazaripouya
  • 04
    Enhancing Operational Grid Resilience Against Wildfires Under Decision-Dependent Uncertainties
    arXiv:2608.02978v1 Announce Type: new Abstract: This paper proposes a new automated decision-making framework to enhance the resilience of electrical systems against wildfires by applying operational strategies that account for decision-dependent uncertainty (DDU). The proposed framework incorporates both preventive and corrective measures, enabling adaptive and automated decision-making throughout the course of evolving wildfire scenarios. First, a baseline multistage optimization model is presArastoo H Salimi, Hamidreza Nazaripouya
  • 05
    Process-Knowledge-Embedded Safe DRL for Real-Time Dispatch of Process Loads in Industrial Microgrids
    arXiv:2608.03149v1 Announce Type: new Abstract: Steelmaking process loads (SPLs) are flexible resources that enhance local renewable-energy utilization and reduce electricity procurement costs in industrial microgrids. However, strong multistage coupling makes current decisions affect subsequent feasibility, challenging conventional deep reinforcement learning to reduce costs while maintaining process feasibility throughout production. This paper proposes a process-knowledge-embedded safe deep rDaniyaer Paizulamu, Lin Cheng, Fashun Shi, Yuchi Zhang, Zhaoyang Dong
  • 06
    Calibration of a Macroscopic Coupled People-Epidemic Transport PDE Model via Density and Velocity Computation from Microscopic Data
    arXiv:2608.03157v1 Announce Type: new Abstract: We introduce an approach for derivation and smoothing of macroscopic densities and velocities from people trajectories, obtained from microscopic position data of individuals within a football stadium. We compute macroscopic densities specifically for susceptible, infected, and exposed individuals, via detection of exposed individuals based on the duration of critical contacts between susceptible and infected. We then present and numerically solveNikoletta Hadjihabi, Nikolaos Bekiaris-Liberis
  • 07
    GriD-LMIA: A Gridding-Based Assembler for Solving Differentiable Parameter-Dependent Linear Matrix Inequalities
    arXiv:2608.03175v1 Announce Type: new Abstract: Parameter-dependent linear matrix inequalities (PD-LMIs) require holding over a continuous domain. When the scheduling parameters vary with time, derivatives of parameter-dependent decisions may also enter the conditions. Since semidefinite programming solvers require finitely many constraints, we introduce GriD-LMIA, the Gridding-based Differentiable PD-LMI Assembler. It converts the conditions that need to hold on a continuous domain into finitelYicheng Xu, Faryar Jabbari
  • 08
    Dynamic Flexibility Requests in Local Flexibility Markets: Quantifying the DSO Willingness to Pay
    arXiv:2608.03226v1 Announce Type: new Abstract: Local Flexibility Markets (LFMs) require Distribution System Operators (DSOs) to determine both the quantity of flexibility to procure and the corresponding willingness to pay during market clearing. Existing approaches typically rely on unrealistic centralized AC-OPF clearing algorithms or strictly localized, static flexibility requests driven primarily by congestion management, while the economic value of flexibility is largely neglected. This paSavvas Panagi, Chrysovalantis Spanias, Petros Aristidou
  • 09
    Partitioned Mixed Small Gain-Phase Decentralized Stability Criterion for Power Systems
    arXiv:2608.03641v1 Announce Type: new Abstract: The increasing penetration of converter-interfaced resources is making power-system stability assessment more challenging, particularly in heterogeneous grids containing both grid-forming and grid-following converters. Existing decentralized mixed small-gain and small-phase criteria provide scalable stability certificates, but they require all converters to satisfy the same type of condition at a given frequency. As a result, they cannot simultaneoDiego Cifelli, Adolfo Anta
  • 10
    Precision Specimen Positioning in Electron Microscopy through Hysteresis Compensation, Iterative Learning, and Vision-Based Sensing
    arXiv:2608.03669v1 Announce Type: new Abstract: Electron microscopy requires nanometer-scale specimen positioning over a long stroke. Piezo-stepper actuators are well suited for this task, but their accuracy is limited by hysteresis, mechanical misalignments, and non-collocated sensing. Prior work has addressed these limitations on simplified lab setups. However, extending to a full electron microscope stage introduces coupled nonlinear kinematics and, importantly, the absence of a dedicated poiJ. S. van Hulst, A. M. C. de Peffer, D. Herceg, E. M. Franken, E. Verschueren, W. P. M. H. Heemels, D. J. Antunes
  • 11
    A Conductance Based Amygdala Model of Threat Processing in Anxiety and Depression
    arXiv:2608.03712v1 Announce Type: new Abstract: Anxiety and depressive disorders are increasingly viewed as dysregulations along continuous stress-regulatory dimensions. However, existing computational approaches seldom connect interpretable circuit level mechanisms to autonomic physiology. Methods: This study develops a mechanistic framework that links amygdala dysregulation to cardiovascular stress responses for digital phenotyping and clinical interpretation. We formulated a compact, nine equMalik Faizan, P. J. White, Indrakshi Dey
  • 12
    Input-to-State Stability of Reset-Integral Sliding Mode Control for Linear Systems
    arXiv:2608.03802v1 Announce Type: new Abstract: This work presents a stability analysis of a hybrid control system integrating a reset controller (RC) featuring a single reset state with an integral sliding-mode controller (ISMC). It is shown that the reachability of the sliding surface is decoupled from the nominal reset mechanism. This decoupling property enables a Lyapunov-based stability analysis, demonstrating that the closed-loop RC-ISMC system achieves input-to-state stability (ISS) and uXinxin Zhang, Leonid Freidovich
  • 13
    A Stackelberg-Bayesian Capacity-Market Game of Carbon Regulation and Second-Life Battery Investment under AI Data-Center Load Growth
    arXiv:2608.03989v1 Announce Type: new Abstract: Artificial intelligence (AI) data centers are driving rapid electricity load growth across all U.S. ISO/RTO regions, raising both system costs and carbon exposure. This study develops a three-level Stackelberg--Bayesian game in which a regulator (leader) sets carbon penalties and subsidies, a single ISO capacity market clears against an energy balance modeled as a classical generation-expansion problem, and technology-specific investors (followers)Rouzbeh Haghighi, Ali Hassan, Sina Mohammadi, Marcus Chen I Wada, Wencong Su
  • 14
    PACE-QAOA: Physics-Constrained Quantum Optimization for Qubit-Efficient Power System Islanding
    arXiv:2608.02789v1 Announce Type: cross Abstract: Controlled islanding partitions a stressed power network to limit disrupted power transfer while preserving operational integrity in every island. This NP-hard partitioning problem becomes increasingly demanding as networks grow, motivating quantum optimization as a complementary approach. However, limited qubit capacity restricts the scale at which conventional QAOA can address islanding. This paper develops a qubit-efficient hybrid quantum formYuqi Jiang, Zhiding Liang, Qiang Guan, Yan Li, Ganesh Kumar Venayagamoorthy
  • 15
    Staying on Spec: Real-Time Monitoring under Uncertainty with a Maritime Case Study
    arXiv:2608.02811v1 Announce Type: cross Abstract: Robotic systems must operate under uncertainty while satisfying complex task and safety specifications. Monitoring such specifications under uncertainty remains challenging, as existing formulations typically require extensive data or explicit uncertainty distributions. In this paper, we propose a real-time monitoring framework that reduces data requirements by leveraging data-driven reachable sets for specification evaluation. We instantiate theElizabeth Dietrich, Hanna Krasowski, Emir Cem Gezer, Roger Skjetne, Asgeir Johan S{\o}rensen, Murat Arcak
  • 16
    Biconvex Optimization for Smooth Minimum-Time Trajectories around Convex Obstacles
    arXiv:2608.02834v1 Announce Type: cross Abstract: We present a biconvex approach for minimum-time motion planning around convex obstacles that is guaranteed to converge, is anytime, and supports derivative constraints to arbitrary order. We jointly convexify the minimum-time objective and all derivative constraints through a change of variables, and handle collision avoidance via time-varying separating planes, reducing the problem to a biconvex program. This program is solved by alternating betPeter Werner, Tobia Marcucci, Daniela Rus
  • 17
    Control Barrier Functions via Minkowski Operations for Safe Navigation among Polytopes
    arXiv:2608.02886v1 Announce Type: cross Abstract: Safely navigating polytopic environments while respecting the dynamics, control, and exact geometry of the underlying system is a challenge in robotics. Control barrier functions (CBFs) synthesize safe control policies by rendering the safe set forward invariant, but many existing CBF-based methods approximate polytopes using conservative smooth shapes, such as spheres or ellipsoids, to obtain explicit differentiable distance functions. In this aYi-Hsuan Chen, Shuo Liu, Wei Xiao, Calin Belta, Michael Otte
  • 18
    A Physics-Informed Hybrid Neural Operator for Transient Magnetization Prediction in Power Magnetics
    arXiv:2608.02965v1 Announce Type: cross Abstract: Magnetic components in high-frequency, high-power-density converters are increasingly driven by non-sinusoidal flux-density waveforms with fast transitions, minor-loop operation, dc bias, and temperature variation. Under these conditions, steady-state core-loss formulas and single-valued material curves cannot fully capture transient magnetization responses. This work proposes the Physics-Informed Hybrid Neural Operator (PI-HNO), a compact materiYachao Zhu, Qiujie Huang, Sinan Li, Yang Li, Gang Lei, Jianguo Zhu
  • 19
    CUDA MPC: A GPU-Native Solver for Model Predictive Control
    arXiv:2608.03051v1 Announce Type: cross Abstract: Model Predictive Control (MPC) delivers constraint-aware control, but its reliance on online optimization limits its use on systems with fast dynamics, high-dimensional models, or long horizons. Existing GPU implementations typically treat the device as a linear-algebra accelerator, leaving the optimization loop dependent on repeated kernel launches and high-latency memory transfers. This paper introduces CUDA MPC, a GPU-native MPC framework thatBabak Akbari, Melissa Greeff
  • 20
    Shooting for Contact: Contact-Implicit Multiple Shooting for Dynamic Motion Retargeting
    arXiv:2608.03116v1 Announce Type: cross Abstract: Motion retargeting approaches often prioritize kinematic similarity over whole-body dynamics, contact consistency, and actuation limits, yielding references that are difficult for reinforcement learning (RL) policies to reproduce, particularly for contact-rich behaviors. We present a contact-implicit, direct simulation-based multiple shooting (DSMS) framework that transforms kinematically feasible references into dynamically feasible whole-body tSergio A. Esteban, Jason H. K. Siu, Derrick Mach, Junheng Li, Vince Kurtz, Joel W. Burdick, Aaron D. Ames
  • 21
    Joint-Range Inequalities for Nonconvex QCQPs
    arXiv:2608.03318v1 Announce Type: cross Abstract: We study cutting planes for nonconvex quadratically constrained quadratic programs (QCQPs) through a project-then-lift approach inspired by mixed-integer rounding (MIR) inequalities. Given two base valid inequalities for the extended QCQP formulation, we project the associated two-row relaxation into a two-dimensional set and analyze the joint range of quadratic functions in two base inequalities. For the nonconvex joint range, we give a closed-fLiding Xu, Sebastian Pokutta
  • 22
    Shaping Wind-Tunnel Airflow for Unmanned Aerial Vehicles using Online Learning
    arXiv:2608.03378v1 Announce Type: cross Abstract: The development and testing of advanced aerial robots require experiments in controlled environments with tailored airflow profiles. This paper presents an online learning algorithm for controlling the complex airflow field in a multi-fan vertical wind tunnel. Our method combines a simplified physical model with iterative, measurement-based learning, enabling sample-efficient convergence to desired airflow distributions. We demonstrate the methodGhadeer Elmkaiel, Michael Muehlebach
  • 23
    Principles of Robot Autonomy
    arXiv:2608.03496v1 Announce Type: cross Abstract: Autonomous robots are moving rapidly from research labs into everyday life - on roads, in the air, in warehouses, and in space. Robot autonomy is no longer solely an academic pursuit, but a collection of mature, field-tested methods and tools that practitioners rely on in real-world deployments. This book offers a clear, unified introduction to the methods that make this possible. Built on decades of teaching at Stanford, the text develops the coDaniele Gammelli, Joseph Lorenzetti, Katie Luo, Gioele Zardini, Marco Pavone
  • 24
    Active Stiffness Control of a Supportive Continuum Robot
    arXiv:2608.03677v1 Announce Type: cross Abstract: Supportive continuum robots (SCRs) enhance the load-bearing capability of an operative continuum robot by mechanically coupling it with a supportive arm. However, their passive stiffness is determined by the mechanical configuration and cannot be adjusted online for varying payloads or interaction forces. Active stiffness control is therefore needed to regulate the load response and maintain positioning accuracy. Meanwhile, the closed-chain strucRana Danesh, Farrokh Janabi-Sharifi, Farhad Aghili
  • 25
    Fidelity-Based Robustness Margins for Finite-Time Quantum Control
    arXiv:2608.03698v1 Announce Type: cross Abstract: We develop a structure-specific fidelity-threshold robustness margin for finite-dimensional closed quantum systems under piecewise-constant coherent control. A scalar physical parameter may perturb the drift, a control Hamiltonian, or another declared Hamiltonian component across the control horizon. A differential sensitivity bound for trace-amplitude gate fidelity yields a threshold-dependent Lipschitz constant on the connected safe parameter cS. P. O'Neil, F. C. Langbein, C. A. Weidner, E. A. Jonckheere, S. Schirmer
  • 26
    ADMITBench: A Safety-Governed Reference Framework for Evaluating the Admissibility of Industrial LLM Advisories
    arXiv:2608.03866v1 Announce Type: cross Abstract: This white paper presents ADMITBench, a reference framework for evaluating industrial LLM advisories at the level of the proposed action. The framework implements a versioned, safety-governed evaluation contract that checks whether a recommendation is supported by the available evidence, permitted under the stated authority and procedure, and acceptable under the plant-specific consequence checks encoded in the selected evaluation profile. In thiYash Misra, Javal Vyas, Siddharth Gutta, Mehmet Mercang\"oz
  • 27
    Operationally Feasible Synthetic Power-Grid Scenarios via Learning the AC-Operable Joint Distribution
    arXiv:2608.03878v1 Announce Type: cross Abstract: Synthetic power-grid scenarios are essential for planning, resilience assessment, contingency analysis, and data-driven power-system applications. Recent synthetic grid generation methods have improved structural realism and operational feasibility by incorporating engineering knowledge through post-generation validation, optimization, or physics-aware generation. However, generated scenarios may still exhibit low AC feasibility and robustness, lChenhan Xiao, Xinyu He, Haoran Li, Hanghang Tong, Yang Weng
  • 28
    Bimanual Manipulation Within an 8 GB Budget: Zero-Copy Sensing and Quantized ACT on an Entry-Level Jetson
    arXiv:2608.03938v1 Announce Type: cross Abstract: Bimanual manipulation policies trained with imitation learning are typically evaluated on workstation or datacenter-class GPUs, leaving the cost of deploying them on embedded hardware largely uncharacterized. We present a bimanual SO-101 system running entirely on an NVIDIA Jetson Orin Nano Super (8 GB), the entry-level tier of NVIDIA's embedded line, using a desktop GPU (RTX 3070) only for offline training, evaluated on pick-and-place of a deforEkansh Singh, Eva Samuel, Alessandra Reneau, Ryan Schmeelk, Yashvi Gandhi
  • 29
    Constrained Performance Boosting Control for Nonlinear Systems
    arXiv:2511.02389v2 Announce Type: replace Abstract: We present the Alternating Direction Method of Multipliers (ADMM) for Performance Boosting (PB), an approach for designing neural controllers for stable nonlinear systems subject to state and input constraints. The method builds on an internal model control formulation of PB. In this setting, the controller is parametrized as a stable neural operator, so closed-loop stability is guaranteed by construction, and its weights are trained offline toGianluca Giacomelli, Danilo Saccani, Siep Weiland, Giancarlo Ferrari-Trecate, Valentina Breschi
  • 30
    Estimating Density Functions for Probabilistic Power Flow Using Invertible Neural Networks
    arXiv:2604.00673v2 Announce Type: replace Abstract: Probabilistic power flow (PPF) is essential for quantifying operational uncertainty in modern power systems with high penetrations of renewable generation and flexible loads. Conventional PPF methods primarily rely on Monte Carlo (MC)- based power flow (PF) simulations or simplified approximations of voltage probability density functions. Although MC methods provide high accuracy, they incur substantial computational and data-storage costs, wheWeijie Xia, James Ciyu Qin, Edgar Mauricio Salazar Duque, Hongjin Du, Peter Palensky, Giovanni Sansavini, Pedro P. Vergara
18天前更新
  • 01
    Detecting high-frequency brain disorder signals using dynamic mode decomposition from EEG
    arXiv:2608.02804v1 Announce Type: new Abstract: Recent studies have reported clearly identifiable dynamical changes in the high-frequency range of EEG signals recorded during specific stimuli, such as visual or auditory inputs, or in cases of brain disorders like epileptic seizures. In this study, we utilized Dynamic Mode Decomposition (DMD) to extract consistent and persistent dynamical changes in the high-frequency band from the signals of neurologically relevant EEG channels. High-frequency DJacob Kang, Jong-Hyeon Seo
  • 02
    Persistent homology broadens the controllable subspace in human structural connectomes
    arXiv:2608.03181v1 Announce Type: new Abstract: Network control theory applied to structural connectomes typically ranks brain regions as candidate driver nodes by their structural connectivity strength, and evaluates performance through scalar control energy. We test whether this framing captures the most relevant information about how driver-node selection shapes brain network control. We introduce an alternative criterion based on the persistent topological cycles in which each node participaCarter Sale, Marco Coraggio, Mengsen Zhang, Michael J. Richardson
  • 03
    Predictive Set Theory: A Generative Framework for Cognitive Architecture with Operationalized Core Mechanisms
    arXiv:2608.02704v1 Announce Type: cross Abstract: Predictive processing theories portray the brain as a hierarchical prediction engine that minimizes prediction error, yet they lack operational definitions for the structure of a "prediction," the standardized response to a prediction error, and the mechanism that maintains consistency across successive updates. Bayesian cognitive science attempts to subsume all uncertainty under probabilistic belief updating, but it presupposes a closed hypothesYiyang Yu
  • 04
    Modelling temporal dynamics of suicidal ideation and behaviour across pre- to early adolescence using a Markov framework
    arXiv:2608.02896v1 Announce Type: cross Abstract: Understanding the dynamics of suicidal ideation and behaviour in youth and the factors associated with transitions from thoughts to behaviours is critical for early identification, monitoring, and prevention. Using longitudinal self-report data from the Adolescent Brain Cognitive Development (ABCD) Study (n = 11,864) spanning ages 9 to 13 years, we developed a time-inhomogeneous discrete-time Markov chain framework to model transitions across eigSieun Lee, Ben Cardoen, Marianne Etherson, Nitish Jawahar, Ellen Townsend, Kapil Sayal, Peter Fonagy, Aja Murray, Joanna Lockwood, Ayan Mahamud, Chris Hollis, Rory O'Connor, Dorothee Auer
  • 05
    A Landau-Ginzburg Phenomenology of Sleep-Stage Transitions
    arXiv:2608.03000v1 Announce Type: cross Abstract: Sleep staging provides a reproducible clinical description, but it does not by itself explain why some boundaries are abrupt while others are graded, or why transition windows contain instability, synchrony, and apparent state coexistence. We develop a local Landau-Ginzburg phenomenology in which each boundary is represented by motion in an effective potential of a spatially extended, noisy, dissipative neural field. A latent cortical-ordering coAlexander Poltorak
  • 06
    MIMIC-MJX: Neuromechanical Emulation of Animal Behavior
    arXiv:2511.20532v3 Announce Type: replace Abstract: The primary output of the nervous system is movement and behavior. While recent advances have democratized pose tracking during complex behavior, kinematic trajectories alone provide only indirect access to the underlying control processes. Here we present MIMIC-MJX, a framework for learning biomechanically grounded neural control policies from kinematics. MIMIC-MJX provides a platform for modeling the generative process of motor control by traCharles Y. Zhang (Harvard University), Yuanjia Yang (Salk Institute for Biological Studies), Aidan Sirbu (Mila), Elliott T. T. Abe (University of Washington), Emil W\"arnberg (Harvard University), Eric J. Leonardis (Salk Institute for Biological Studies), Diego E. Aldarondo (Harvard University), Adam Lee (Harvard University), Aaditya Prasad (Massachusetts Institute of Technology), Jason Foat (Salk Institute for Biological Studies), Kaiwen Bian (Salk Institute for Biological Studies), Joshua Park (Salk Institute for Biological Studies), Rusham Bhatt (Salk Institute for Biological Studies), Vyom N. Patel (Neuromatch), Hutton Saunders (Salk Institute for Biological Studies), Austin O. Barbano (Salk Institute for Biological Studies), Akira Nagamori (Salk Institute for Biological Studies), Ayesha R. Thanawalla (Salk Institute for Biological Studies), Kee Wui Huang (Salk Institute for Biological Studies), Fabian Plum (Imperial College London), Hendrik K. Beck (Imperial College London), Steven W. Flavell (Massachusetts Institute of Technology), David Labonte (Imperial College London), Blake A. Richards (Mila), Bingni W. Brunton (University of Washington), Eiman Azim (Salk Institute for Biological Studies), Bence P. \"Olveczky (Harvard University), Talmo D. Pereira (Salk Institute for Biological Studies)
  • 07
    Do VLMs Align Better with Humans than LLMs during Natural Reading?
    arXiv:2605.28818v2 Announce Type: replace-cross Abstract: Large language models have become increasingly useful computational models of human language processing, but it remains open whether vision-language learning makes text representations more human-like during natural reading. We address this question by comparing matched LLM and vision-language model pairs under strictly text-only input and evaluating alignment with human brain activity (whole-cortex fMRI) and human behavior (synchronizedJinzhou Wu, Zhengwu Ma, Jixing Li, Baoping Tang, Zitong Lu
cs.NE
arXiv
18天前更新
  • 01
    NeuroMosaic: Anatomically Grounded Multimodal Large Language Modeling for Molecularly Aware Glioma Reasoning from 3D MRI and Clinical Narratives
    arXiv:2608.03187v1 Announce Type: new Abstract: Multimodal medical large language models remain structurally weak for neuro-oncology because volumetric evidence is compressed into generic visual tokens and diagnostic conclusions often lack an auditable link to MRI regions. We present NeuroMosaic, a 3D multimodal language model that converts multi-sequence brain MRI into anatomy-indexed regional tokens, aligns them with clinical narrative and molecular concepts, and generates evidence-linked outpYantong Liu, Zheyu Zhang, Runpeng Liu, Mu Xitang, Seong-Yoon Shin, Hyun-Ae Lee
  • 02
    Impacts of Single-objective Landscapes on Multi-objective Optimization
    arXiv:2608.03266v1 Announce Type: new Abstract: This work revealed a relationship between a multi-objective optimization problem and single-objective optimization problems that exist in the multi-objective problem. This work focused on combinatorial problems and investigated the relations between the local optima networks of the single-objective problems and the Pareto optima network of the multi-objective problem. Each of their networks has a graph structure. We divided the entire network intoShoichiro Tanaka, Keiki Takadama, Hiroyuki Sato
  • 03
    MuEvo: LLM-Driven Evolution of Multi-Heuristic Ensemble
    arXiv:2608.03636v1 Announce Type: new Abstract: Large language model-based automated heuristic design (LLM-AHD) has shown strong potential in discovering effective heuristics for combinatorial optimization problems. However, existing methods primarily optimize a single heuristic, whereas practical optimization frameworks often rely on multiple interacting components. Directly extending single-heuristic methods is challenging because early component selection can overlook components with late potHaoze Lv, Ning Lu, Shengcai Liu, Shaofeng Zhang, Ke Tang
  • 04
    Self-Organising Digital Circuits
    arXiv:2608.02606v1 Announce Type: cross Abstract: Fault tolerance in classical computing has traditionally relied on static strategies like hardware redundancy and error-correcting codes. Biological systems, in contrast, exhibit adaptive plasticity, maintaining function through dynamic re-organisation around damage. Inspired by this principle, we introduce Self-Organising Digital Circuits, framing functional logic generation and maintenance as a meta-learning problem on graphs. Our architectureMarcello Barylli, Gabriel B\'ena, Alexander Mordvintsev, Eleni Nisioti, Sebastian Risi
  • 05
    BBOWP-Bench: Evaluating LLMs on Black-Box Optimization Word Problems
    arXiv:2608.02612v1 Announce Type: cross Abstract: Formulating an optimization problem strongly affects the quality of the final solution, yet good formulations usually require substantial expertise. Recent studies have therefore examined how to automatically derive optimization problems from natural-language descriptions, but existing benchmarks focus on settings where objectives and constraints can be written explicitly as mathematical expressions. Many practically important problems are naturaYutaro Yamada, Kei Hiroshima, Nozomu Yoshinari, Kento Uchida, Shinichi Shirakawa
  • 06
    AS-FedBridge: Pseudo-Spike Bridge Distillation for Heterogeneous ANN-SNN Federated Learning
    arXiv:2608.03324v1 Announce Type: cross Abstract: Federated learning enables collaborative model training across distributed edge devices while strictly preserving data privacy. To facilitate practical deployment on resource-constrained edge devices, Spiking Neural Networks (SNNs) have emerged as a promising alternative to traditional Artificial Neural Networks (ANNs) due to their sparse computing mechanisms and high energy efficiency. However, jointly training ANNs and SNNs exposes a challengeShengyang Li, Yiting Dong, Liuyang Song, Ximing Wang, Luyuan Xie, Cong Li, Qingni Shen, Zhaofei Yu
  • 07
    Omega-S: A Functional Resilience Index for LLM Fine-Tuning
    arXiv:2608.03887v1 Announce Type: cross Abstract: Fine-tuning a large language model on new data degrades what it previously learned. We present Omega-S, a drop-in penalty computed from the weight matrix alone: it needs no previous-task data, no Fisher matrix and no stored copy of the old weights. It is three lines in an existing training loop and adds under 4% to the cost of a step. Retention. On Llama-3-8B with LoRA, fine-tuned from code to prose and measured by HumanEval over ten seeds, OmegaAlberto Acedo
  • 08
    The Transformer Revolution, Part 1: Dynamic Processing through Output- Weight Interconnections
    arXiv:2608.03921v1 Announce Type: cross Abstract: This paper offers a new interpretation of the Transformer during inference. Against the "stochastic parrot" view that large language models merely reproduce statistical regularities learned in training, we argue that Transformers construct and apply prompt-dependent transformations whose parameters are generated during inference. We call this form of computation SIDPP: Sequence-level Interactive Dynamic Parallel Processing. The Transformer is intMarco Giunti, Fabrizia Giulia Garavaglia
  • 09
    Quantization Effects of Artificial Neural Networks for Embedded Edge-Computing Applications
    arXiv:2511.05479v4 Announce Type: replace Abstract: This paper examines the use of Quantized Neural Networks (QNNs) for two resource-constrained scientific applications: automated calibration of semi-conductor quantum bits (qubits) and scientific particle detectors. We evaluate the trade-offs between Post-Training Quantization (PTQ), Quantization-Aware Training (QAT), and ultra-low-bit Binary Neural Networks (BNNs) with respect to latency and resource usage. Our results demonstrate that PTQ achiAlperen Aksoy, Ilja Bekman, Vesselin Dimitrov, Qader Dorosti, Chimezie Eguzo, Sarah Fleitmann, Christian Grewing, Fabian Hader, Andre Zambanini, Stefan van Waasen
  • 10
    OR-Agent: Bridging Evolutionary Search and Structured Research for Automated Algorithm Discovery
    arXiv:2602.13769v3 Announce Type: replace-cross Abstract: Automating heuristic design in complex, experiment-driven domains requires more than iterative mutation of solution algorithms. Current LLM-based evolutionary methods often rely on stochastic mutation loops that lack long-term strategic planning and a formal mechanism to learn from historical failures, leading to inefficient exploration and redundant trials. To address this, we present OR-Agent, a multi-agent research framework designed fQi Liu, Ruochen Hao, Can Li, Wanjing Ma