全部/知识/实时热榜

OpenReview · 实时热榜

HISTORY2026年8月10日55 不同热搜
08/0309/01 有历史数据
DAILY UNIQUE TOPICS55 个热搜
  1. 01
    The 2025 Foundation Model Transparency Index

    期刊:Transactions on Machine Learning Research · 摘要:Foundation model developers are among the world’s most important companies. As these companies become increasingly consequential, how do their transparency practices evolve? The 2025 Foundation Model Transparency Index is the third edition of an annual effort to characterize and quantify the transparency of foundation model developers. The 2025 FMTI introduces new indicators related to data acquisition, usage data, and monitoring and evaluates companies like Alibaba, DeepSeek, and xAI for the first time. The 2024 FMTI reported that transparency was improving, but the 2025 FMTI finds this prog… · 篇幅:Long submission (more than 12 pages of main content) · OpenReview ID:1jT253Xtyf

    最高第 2100:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  2. 02
    XFeat Revisited: Reproducibility and Evaluation of a Lightweight Image Matcher

    期刊:Transactions on Machine Learning Research · 摘要:We present a reproducibility study of XFeat, a lightweight local feature extractor and matcher designed to identify corresponding points across images efficiently on resource-constrained hardware. We re-implement the architecture based on the paper and supplementary material, re-evaluate the authors' released checkpoint alongside our re-implementation, and conduct additional architectural ablations to examine design choices that were not fully justified in the original work. This distinction between re-evaluation and reproduction is important, as the paper, supplement, and public code differ… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/GalaxyGHz/xfeat-revisited · OpenReview ID:2WI889Ulin

    最高第 1500:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  3. 03
    Reasoning with Preference Constraints: A Benchmark for Language Models in Many-to-One Matching Markets

    期刊:Transactions on Machine Learning Research · 摘要:Recent advances in reasoning with large language models (LLMs) have demonstrated strong performance on complex mathematical tasks. Techniques such as Chain-of-Thought and In-Context Learning have further enhanced this capability, making LLMs both powerful and accessible tools for a wide range of users, including non-experts. However, their application to problems arising in operations research, particularly those at the intersection of combinatorial optimization and game theory that require domain expertise, remains underexplored. To address this gap, we introduce a benchmark of 369 instances… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/maryloufauchard/CAP_Benchmark · OpenReview ID:2dpt2Ughzt

    最高第 4300:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  4. 04
    Enhance the Safety in Reinforcement Learning by ADRC Lagrangian Methods

    期刊:Transactions on Machine Learning Research · 摘要:Safe reinforcement learning (Safe RL) seeks to maximize rewards while satisfying safety constraints, typically addressed through Lagrangian-based methods. However, existing approaches, including PID and classical Lagrangian methods, suffer from oscillations and frequent safety violations due to parameter sensitivity and inherent phase lag. To address these limitations, we propose ADRC-Lagrangian methods that leverage Active Disturbance Rejection Control (ADRC) for enhanced robustness and reduced oscillations. Our unified framework subsumes a broad class of PID Lagrangian updates as frozen-par… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:3IbuT8uzYS

    最高第 1900:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  5. 05
    VQEL: Enabling Self-Play in Emergent Language Games via Agent Internal Vector Quantization

    期刊:Transactions on Machine Learning Research · 摘要:Emergent Language (EL) focuses on the emergence of communication among artificial agents. Although symbolic communication channels more closely mirror the discrete nature of human language, learning such protocols remains fundamentally difficult due to the non-differentiability of symbol sampling. Existing approaches typically rely on high-variance gradient estimators such as REINFORCE or on continuous relaxations such as Gumbel–Softmax, both of which suffer from limitations in training stability and scalability when learning a language from scratch. Motivated by cognitive theories that empha… · 篇幅:Long submission (more than 12 pages of main content) · OpenReview ID:5nqQlGWlsW

    最高第 1700:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  6. 06
    Influencing Humans to Conform to Preference Models for RLHF

    期刊:Transactions on Machine Learning Research · 摘要:Designing a reinforcement learning from human feedback (RLHF) algorithm to approximate a human's unobservable reward function requires assuming, implicitly or explicitly, a model of human preferences. In sequential decision making tasks, a preference model that poorly describes how humans generate preferences risks learning a poor approximation of the human’s reward function. In this paper, we conduct human studies to assess whether one can influence the expression of real human preferences to more closely conform to a desired preference model. Importantly, our approach does not seek to alter… · 篇幅:Long submission (more than 12 pages of main content) · OpenReview ID:7YPlw1nUmW

    最高第 2500:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  7. 07
    Dynamic Reward Incentives for Emergent Cooperation under Changing Rewards

    期刊:Transactions on Machine Learning Research · 摘要:Peer incentivization (PI) is a popular multi-agent reinforcement learning approach where all agents can reward or penalize each other to achieve cooperation in social dilemmas. Despite their potential for scalable cooperation, current PI methods heavily depend on fixed incentive values that need to be appropriately chosen with respect to the environmental rewards and thus are highly sensitive to their changes. Therefore, they fail to maintain cooperation under changing rewards in the environment, e.g., caused by modified specifications, varying supply and demand, or sensory flaws — even when… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/philippaltmann/DRIVE · OpenReview ID:9Ltu1HV2YI

    最高第 4100:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  8. 08
    From Centerlines to Hemodynamics: Anisotropic RBF Decoders for Coronary Arteries

    期刊:Transactions on Machine Learning Research · 摘要:Accurate and rapid estimation of hemodynamic metrics, such as pressure and wall shear stress (WSS), is important for assessing the severity of Coronary Artery Disease (CAD). Existing approaches, including invasive Fractional Flow Reserve (FFR) measurements and computationally expensive Computational Fluid Dynamics (CFD) simulations, face challenges in invasiveness, cost, and speed. We present a learned surrogate for fast prediction of CFD-simulated coronary hemodynamics from vessel centerline geometry. The model encodes 1D vessel centerlines together with inlet flow rate using a transformer-b… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:AoJUrVjufP

    最高第 600:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  9. 09
    On the Fundamental Limits of LLMs at Scale

    期刊:Transactions on Machine Learning Research · 摘要:Large Language Models (LLMs) have benefited enormously from scaling, yet these gains are bounded by five fundamental limitations: (1) hallucination, (2) context compression, (3) reasoning degradation, (4) retrieval fragility, and (5) multimodal misalignment. While existing surveys describe these phenomena empirically, they lack a rigorous theoretical synthesis connecting them to the foundational limits of computation, information, and learning. This work closes that gap by presenting a unified, proof-informed framework that formalizes the innate theoretical ceilings of LLM scaling. First, com… · 篇幅:Long submission (more than 12 pages of main content) · OpenReview ID:BIRDGVrom8

    最高第 3700:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  10. 10
    NeMoS: Nearest Neighbors Bandit meets Active Learning for Online Model Selection

    期刊:Transactions on Machine Learning Research · 摘要:The proliferation of open-platform text-to-image generative models has made prompt-wise model selection critical to maximize generation quality and semantic alignment. However, current strategies, such as contextual bandits, often converge slowly and fail to exploit the semantic relationships across prompts. To bridge this gap, we propose NeMoS, a non-parametric bandit framework that couples nearest neighbor reward estimation with a budget-constrained active learning strategy. Specifically, our approach operates in the prompt embedding space and estimates the reward of incoming prompts based… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/julesdamidaux/nemos-tmlr · OpenReview ID:CSjewjplO1

    最高第 4500:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  11. 11
    PAC-Bayesian Meta-Learning for Few-Shot Identification of Linear Dynamical Systems

    期刊:Transactions on Machine Learning Research · 摘要:Identifying linear time-invariant (LTI) dynamical systems from data is especially challenging when trajectories are short, noisy, or high-dimensional. Traditional system identification methods typically treat each system in isolation and therefore fail to exploit shared structure across related systems. We propose a PAC-Bayesian meta-learning framework for few-shot LTI system identification (PBML-LTI), which learns a transferable prior over task-specific dynamics while preserving task-level heterogeneity. Each task corresponds to an unknown LTI system, and a meta-learner uses a collection of… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/chenfeng-huang/PBML-LTI-TMLR-2026 · OpenReview ID:CiGFpSLzFv

    最高第 4200:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  12. 12
    Toward Greater Autonomy in Materials Discovery Agents: Unifying Planning, Physics, and Scientists

    期刊:Transactions on Machine Learning Research · 摘要:We aim at designing language agents with greater autonomy for crystal materials discovery. While most of existing studies restrict the agents to perform specific tasks within predefined workflows, we aim to automate workflow planning given high-level goals and scientist intuition. To this end, we propose Materials Agent unifying Planning, Physics, and Scientists, known as MAPPS. MAPPS consists of a Workflow Planner, a Tool Code Generator, and a Scientific Mediator. The Workflow Planner uses large language models (LLMs) to generate structured and multi-step workflows. The Tool Code Generator s… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:Cwq1U8tbWW

    最高第 800:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  13. 13
    Optimal or Greedy Decision Trees? Revisiting their Objectives, Tuning, and Performance

    期刊:Transactions on Machine Learning Research · 摘要:Recently there has been a surge of interest in optimal decision tree (ODT) methods that globally optimize accuracy directly, in contrast to traditional approaches that locally optimize an impurity or information metric. However, the literature shows conflicting evidence on the value of ODTs, with some demonstrating superior out-of-sample performance of ODTs over greedy approaches, while others show the opposite. The value and performance of ODTs therefore remains one of several open question regarding ODTs, most of which could not be answered before due to lack of scalability. With our experi… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/ConSol-Lab/opt-vs-greedy-dts · OpenReview ID:DvDOAtskXl

    最高第 2200:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  14. 14
    Predicting Chain-of-Thought Correctness from Trajectory Geometry

    期刊:Transactions on Machine Learning Research · 摘要:We ask whether the geometry of a reasoning trajectory, that is, how a chain-of-thought (CoT) trace moves through semantic space beyond its raw length, predicts whether the final answer is correct, and whether that prediction is useful in practice. Across 2,800 CoT traces spanning three reasoning benchmarks (FOLIO, GSM8K, and PrOntoQA) and five language models, we extract interpretable trajectory-level features (adjacent-step transition energy, path entropy, semantic drift, loopiness, discourse-graph spectra, and direction-sensitive drift) and predict per-trace correctness. Under problem-group… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:H9cBkEqVeY

    最高第 900:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  15. 15
    Gen-MURE: Generalized Multiplicative Unbiased Risk Estimate

    期刊:Transactions on Machine Learning Research · 摘要:Coherent imaging modalities such as ultrasound and synthetic aperture radar (SAR) images are degraded by signal-dependent multiplicative noise, where the noise distributions vary widely across acquisition scenarios. Existing self-supervised image denoising methods either assume zero-mean additive noise, independence across pixels or require the noise distribution to be known, which often limit their applicability in real-world image denoising systems. We propose a Generalized Multiplicative Unbiased Risk Estimate (Gen-MURE), a model-agnostic self-supervised image denoising framework for enhan… · 篇幅:Long submission (more than 12 pages of main content) · OpenReview ID:Hie13qRm1x

    最高第 1000:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  16. 16
    Probing Visual Concepts in Lightweight Vision-Language Models for Automated Driving

    期刊:Transactions on Machine Learning Research · 摘要:The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning and generalisation capabilities to handle long-tail scenarios. However, these models often fail on simple visual questions that are highly relevant to automated driving, and the reasons behind these failures remain poorly understood. In this work, we examine the intermediate activations of VLMs and assess the extent to which specific visual concepts are linearly encoded, with the goal of identifying bottlenecks in the flow of visual information… · 篇幅:Long submission (more than 12 pages of main content) · OpenReview ID:HlBBy19ojC

    最高第 3000:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  17. 17
    Clarity: The Flexibility-Interpretability Trade-Off in Sparsity-aware Concept Bottleneck Models

    期刊:Transactions on Machine Learning Research · 摘要:The widespread adoption of deep learning models in computer vision has intensified concerns about interpretability. Despite strong performance, these models are often treated as black boxes, with limited systematic investigation of their decision-making processes. While many interpretability methods exist, objective evaluation of learned representations remains limited, particularly for approaches that rely on sparsity to ``induce'' interpretability. In this work, we investigate how modeling choices in Concept Bottleneck Models (CBMs) affect the semantic alignment of concept representations.… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/konpanousis/cbm-clarity · OpenReview ID:IyQEQBRR4M

    最高第 300:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  18. 18
    Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning

    期刊:Transactions on Machine Learning Research · 摘要:Learning general-purpose reasoning capabilities has long been a challenging problem in AI. Recent research in LLMs, such as DeepSeek-R1, has shown that reinforcement learning techniques like GRPO enable pre-trained LLMs to develop reasoning capabilities using simple question-answer pairs. In this paper, we aim to train visual language models (VLMs) to perform reasoning on image data through reinforcement learning and visual question-answer pairs, without explicitly using any chain-of-thought (CoT) supervision. Our key finding indicates that simply applying GRPO to a VLM---by prompting the mod… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/maifoundations/Visionary-R1 · OpenReview ID:JWkZXBgh5a

    最高第 4400:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  19. 19
    A Cross-Model Study of Over-Compliance in Large Lan- guage Models

    期刊:Transactions on Machine Learning Research · 摘要:Large language models increasingly mediate decisions in healthcare, legal advisory, and financial analysis, settings in which a model’s willingness to answer an inadequate prompt can matter as much as the accuracy of its answer. Yet systematic cross-model evidence on this behavior remains scarce. The present study examined over-compliance, understood as the generation of substantive content when the input warrants clarification, refusal, or deferral. Four frontier models from Ope- nAI, Google, Meta, and Anthropic were evaluated on a benchmark of 400 prompts spanning under- specification, ambi… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/hgus107/LLM-Over-Complaince · OpenReview ID:LnUP74YNze

    最高第 1100:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  20. 20
    How Much Information Fits in a Vector?

    期刊:Transactions on Machine Learning Research · 摘要:Recent work in neural network interpretability has suggested that hidden activations of some deep models can be viewed as linear projections of much higher-dimensional vectors of sparse latent ``features.'' In general, this kind of representation is known as a superposition code. This work presents an information-theoretic account of superposition codes in a setting applicable to interpretability. We show that when the number $k$ of active features is very small compared to the number $N$ of total features, simple inference methods currently used by sparse autoencoders can reliably decode a $… · 篇幅:Long submission (more than 12 pages of main content) · OpenReview ID:Nby4pCPIZI

    最高第 3400:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  21. 21
    Analyzing the Effect of Noise in LLM Fine-Tuning

    期刊:Transactions on Machine Learning Research · 摘要:Fine-Tuning is the dominant paradigm for adapting pretrained large language models (LLMs) to downstream NLP tasks. In practice, fine-tuning datasets may contain various forms of noise that arise from annotation errors or automated data collection. Although prior work has concentrated on designing robust learning algorithms to mitigate performance degradation under noisy conditions, comparatively little is known about how different types of noise affect the internal learning dynamics of LLMs during fine-tuning. In this work, we systematically study the impact of noise on model behaviour across… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/LingfangLi/analyzing-noise-llm-finetuning · OpenReview ID:NlSBeHZEz5

    最高第 400:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  22. 22
    SAFT: Structure-Aware Fine-Tuning of Large Language Models for AMR-to-Text Generation

    期刊:Transactions on Machine Learning Research · 摘要:Large Language Models (LLMs) are increasingly applied to tasks involving structured inputs such as semantic graphs, yet adapting them to such inputs remains non-trivial. Common approaches either linearize graphs, discarding structural information, or rely on specialized architectures that are not directly compatible with standard pretrained LLMs. We present SAFT, a structure-aware fine-tuning method that augments LLMs with graph positional encodings derived from the magnetic Laplacian of the input graph. These encodings are projected into the LLM embedding space, introducing relational induct… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/guerrantif/saft · OpenReview ID:QZoUMyzYDB

    最高第 2300:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  23. 23
    eDQA: Efficient Deep Quantization of DNN Activations on Edge Devices

    期刊:Transactions on Machine Learning Research · 摘要:Quantization of Deep Neural Network (DNN) activations is a commonly used technique to reduce compute and memory demands during DNN inference, which can be particularly beneficial on resource-constrained edge devices. To achieve high accuracy, existing methods for quantizing activations rely on complex mathematical computations or perform extensive online searches for the best hyperparameters. However, these expensive operations are impractical on edge devices with limited computational capabilities, memory capacities, and energy budgets. Furthermore, many existing methods either do not focus… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/gicLAB/eDQA · OpenReview ID:SEIBCdgE5W

    最高第 1600:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  24. 24
    Learning Structured Set Utility Functions with Contrastive Element Representations

    期刊:Transactions on Machine Learning Research · 摘要:Learning utility functions over sets of elements is central to many machine learning and decision-making tasks such as feature selection, sensor placement, and content recommendation, where the goal is to evaluate and select an optimal subset of elements that provide the largest utility. These utility functions often exhibit desirable properties like monotonicity and submodularity over sets, but are typically expensive to evaluate and may lack an explicit analytical form. Moreover, the utility of a set can vary depending on certain contextual variables, further complicating the learning task.… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:SZ8mOziJBx

    最高第 3200:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  25. 25
    Instance-Level Generation for Representation Learning

    期刊:Transactions on Machine Learning Research · 摘要:Instance-level recognition (ILR) focuses on identifying individual objects rather than broad categories, offering the highest granularity in image classification. However, this fine-grained nature makes creating large-scale annotated datasets challenging, limiting ILR’s real-world applicability across domains. To overcome this, we introduce a novel approach that synthetically generates diverse object instances from multiple domains under varied conditions and backgrounds, forming a large-scale training set. Unlike prior work on automatic data synthesis, our method is the first to address ILR-… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/yankungou/ILGen · OpenReview ID:T3JgJXH3ZK

    最高第 3100:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  26. 26
    The Impact of Enforcing Representational Consistency of Identical Transformations for Disentangled Representation

    期刊:Transactions on Machine Learning Research · 摘要:Recent symmetry-based approaches in Variational Autoencoders (VAEs) have advanced disentanglement learning and compositional generalization. However, existing methods can encode identical semantic transformations differently depending on the specific sample pairs, which reduce the representational consistency of identical transformations. In this paper, we analyze how three commonly used symmetry parameterization families in prior work, namely (1) matrix-exponential parameterizations over the general linear group GL(n), (2) vector-additive actions in latent space, and (3) surjective mappings… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/GIST-IRR/RCIT · OpenReview ID:VjbBxj4aWb

    最高第 2900:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  27. 27
    A Survey of Agent Memory in the Second Half: Towards Self-Evolving and Long-Horizon Agents

    期刊:Transactions on Machine Learning Research · 摘要:Research in artificial intelligence is undergoing a paradigm shift from prioritizing model innovations and benchmark scores towards emphasizing problem definition and rigorous real-world evaluation. As the field enters the "second half," the central challenge becomes real utility in long-horizon, dynamic, and user-dependent settings such as agentic coding, deep research, and computer use, where LLM-based agents face context explosion beyond fixed context windows and must continuously accumulate, manage, and selectively reuse large volumes of information across extended interactions. Memory, w… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/AgentMemoryWorld/Awesome-Agent-Memory · OpenReview ID:XycbogUAeJ

    最高第 3300:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  28. 28
    DS-STAR: Data Science Agent for Solving Diverse Tasks across Heterogeneous Formats and Open-Ended Queries

    期刊:Transactions on Machine Learning Research · 摘要:While large language models (LLMs) have shown promise in automating data science, existing agents often struggle with the complexity of real-world workflows that require exploring multiple sources and synthesizing open-ended insights. In this paper, we introduce DS-STAR, a specialized agent to bridge this gap. Unlike prior approaches, DS-STAR is designed to (1) seamlessly process and integrate data across diverse, heterogeneous formats, and (2) move beyond simple QA to generate comprehensive research reports for open-ended queries. Extensive evaluation shows that DS-STAR achieves state-of-the… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/google-research/ds-star · OpenReview ID:Yz3ZPLzYaU

    最高第 2000:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  29. 29
    Minimax learning rates for estimating binary classifiers under margin conditions

    期刊:Transactions on Machine Learning Research · 摘要:We study classification problems using binary estimators where the decision boundary is described by horizon functions and where the data distribution satisfies a geometric margin condition. A key novelty of our work is the derivation of lower bounds for the worst-case learning rates over broad classes of functions, under a geometric margin condition---a setting that remains theoretically challenging. Moreover, we work in the noiseless setting, where lower bounds are particularly hard to establish. Our general results cover, in particular, classification problems with decision boundaries belo… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:ZIshsqojB6

    最高第 2400:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  30. 30
    Inference-Time Computations for LLM Reasoning and Planning: A Benchmark and Insights

    期刊:Transactions on Machine Learning Research · 摘要:We examine the reasoning and planning capabilities of large language models (LLMs) in solving complex tasks. Recent advances in inference-time techniques demonstrate the potential to enhance LLM reasoning without additional training by exploring intermediate steps during inference. Here, we explore how scaling inference-time techniques can improve reasoning and planning, focusing on understanding the tradeoff between computational cost and performance. To this end, we construct a comprehensive benchmark, known as *Sys2Bench*, and perform extensive experiments evaluating existing inference-tim… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/divelab/sys2bench · OpenReview ID:budZJyCK8G

    最高第 3600:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  31. 31
    Hyperedge Anomaly Detection with Hypergraph Neural Network

    期刊:Transactions on Machine Learning Research · 摘要:Hypergraph is a data structure that enables us to model higher-order associations among data entities. Conventional graph-structured data can represent pairwise relationships only, whereas hypergraph enables us to associate any number of entities, which is essential in many real-life applications. Hypergraph learning algorithms have been well-studied for numerous problem settings, such as node classification, link prediction, etc. However, much less research has been conducted on anomaly detection from hypergraphs. Anomaly detection identifies events that deviate from the usual pattern and ca… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:etPYIk1BqO

    最高第 1300:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  32. 32
    Same-Expert Iteration Improves a Translation MoE Where Expert Communication Does Not

    期刊:Transactions on Machine Learning Research · 摘要:Mixture-of-Experts (MoE) models achieve scalability through sparse expert routing, but experts process tokens independently. A natural hypothesis is that enabling expert communication—through learned topologies, message passing, or sequential chains—should improve performance. We test this hypothesis on WMT14 En-De translation with a small decoder-only Transformer, evaluating ten communication approaches across seven exper-imental axes. We find no clear evidence that any variant improves over standard MoE, though modest sample sizes limit power for detecting small effects; several variants de… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:hPD4MjMfoN

    最高第 500:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  33. 33
    Federated Class-Incremental Learning with Hierarchical Generative Prototypes

    期刊:Transactions on Machine Learning Research · 摘要:Federated Learning (FL) aims at unburdening the training of deep models by distributing computation across multiple devices (clients) while safeguarding data privacy. On top of that, Federated Continual Learning (FCL) also accounts for data distribution evolving over time, mirroring the dynamic nature of real-world environments. While previous studies have identified Catastrophic Forgetting and Client Drift as primary causes of performance degradation in FCL, we shed light on the importance of Incremental Bias and Federated Bias, which cause models to prioritize classes that are recently intr… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/aimagelab/fed-mammoth · OpenReview ID:k2TT42Ei8W

    最高第 4000:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  34. 34
    Analysis of Natural Actor-Critic with Randomized Low- Discrepancy Sampling

    期刊:Transactions on Machine Learning Research · 摘要:Natural gradient methods are appealing in policy optimization due to their invariance to smooth reparameterization and their ability to account for the local geometry of the policy manifold. These properties often lead to improved conditioning of the optimization problem compared to Euclidean policy gradients. However, their reliance on Monte Carlo estimation introduces high variance and sensitivity to hyperparameters. In this paper, we address these limitations by integrating Randomized Quasi-Monte Carlo (RQMC) sampling into the natural actor-critic (NAC) framework. We revisit the NAC linear… · 篇幅:Long submission (more than 12 pages of main content) · OpenReview ID:kOSx9v6dfb

    最高第 2700:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  35. 35
    DIMENSION DOMAIN CO-DECOMPOSITION: SOLVING PDES WITH INTERPRETABILITY

    期刊:Transactions on Machine Learning Research · 摘要:Physics-informed neural networks (PINNs) have demonstrated effectiveness in solving partial differential equations (PDEs), yet they often struggle in high-dimensional regimes and lack interpretable representations and in scenarios involving sharp solution structures. Moreover, existing approaches typically rely on manually specified domain partitions. We propose a unified Dimension–Domain Co-Decomposition (3D) framework that jointly integrates dimension-wise decomposition with mixture-of-experts (MoE)–based domain decomposition. At the dimension level, we introduce an interpretable decomposit… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/TIML-Group/3DPINN · OpenReview ID:kuzkynVyRq

    最高第 3900:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  36. 36
    QMoE+: Hybrid Quantum Mixture of Experts

    期刊:Transactions on Machine Learning Research · 摘要:Quantum mixture of experts (QMoE) extends conditional computation to the NISQ setting by distributing learning across parameterized quantum circuit (PQC) experts selected via a routing mechanism. Existing approaches are limited by single-block experts, lack of load balancing, and aggregation schemes that ignore routing amplitudes. We propose QMoE+, which uses two-block data re-uploading experts with learnable offsets, a coherent aggregation circuit over the joint routing-data Hilbert space, and a Switch-style load-balancing loss. Under top-k=1 sparse routing, QMoE+ activates only ∼28% of its… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/HbjNiser/qmoe-plus · OpenReview ID:l1JaPqZ6K5

    最高第 1800:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  37. 37
    ANU-RL: A New Perspective on Weakly-Supervised Representation Learning for Visual Place Recognition

    期刊:Transactions on Machine Learning Research · 摘要:Representation Learning (RL) is fundamental for image matching, retrieval, classification, and other applications, enabling task-specific feature learning. RL algorithms aim to learn compact embeddings that preserve the neighbourhood structure of the input data. A general approach to this is contrastive learning, which pulls similar images (positives) closer together and pushes dissimilar images (negatives) farther apart in the embedding space. In Visual Place Recognition (VPR), positive images of a query share specific geographical and visual attributes with the query and can form a cluster.… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/Anuradha-Uggi/ANU-RL · OpenReview ID:mXE4OP55il

    最高第 3800:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  38. 38
    Navigating the Labyrinth: Evaluating LLMs’ Ability to Reason About Search Problems

    期刊:Transactions on Machine Learning Research · 摘要:Large Language Models (LLMs) have recently achieved impressive performance in math and reasoning benchmarks. However, they often struggle with logic problems and puzzles that are relatively easy for humans. To further investigate this, we introduce a new benchmark, SearchBench, which contains 11 unique search problems inspired by intuitive puzzles. Each SearchBench problem type is equipped with automated pipelines to generate an arbitrary number of instances and analyze the feasibility, correctness, and optimality of LLM-generated solutions. We show that using step-by-step, language-only reas… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:oub2I1ioL5

    最高第 1200:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  39. 39
    On the Statistical Limits of Self-Improving Agents

    期刊:Transactions on Machine Learning Research · 摘要:We develop a learning-theoretic framework for analyzing self-improving agents by decomposing self-modification into five axes. Within this framework, we prove a sharp boundary: under standard i.i.d. assumptions, distribution-free PAC learnability is preserved if and only if the policy-reachable family remains uniformly capacity-bounded. If reachable capacity can grow without bound, utility-rational self-changes can make learnable tasks unlearnable. We further introduce a simple Two-Gate guardrail—a validation-improvement requirement plus a capacity cap—that preserves this boundary and yields… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:q4vuDMtYgF

    最高第 200:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  40. 40
    Physics-Aware Variational Autoencoder for Urban Travel Demand Calibration

    期刊:Transactions on Machine Learning Research · 摘要:Urban mobility digital twins are revolutionizing how cities manage increasingly complex transportation systems, enabling real-time optimization across multiple stakeholders, services, and dynamic operations. Central to these digital twins is the origin-destination (OD) calibration problem—estimating travel demand patterns that produce realistic traffic simulations matching observed conditions. However, existing calibration methods face critical limitations: they require a prohibitively large number of expensive simulation runs and struggle with high-dimensional city-scale networks. To mitigat… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:r5oS1XXbT3

    最高第 100:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  41. 41
    Environmental Footprint of GenAI Research: Insights from the Moshi Foundation Model

    期刊:Transactions on Machine Learning Research · 摘要:New multi-modal large language models (MLLMs) are continuously being trained and deployed, following rapid development cycles. This generative AI frenzy is driving steady increases in energy consumption, greenhouse gas emissions, and a plethora of other environmental impacts linked to datacenter construction and hardware manufacturing. Mitigating the environmental consequences of GenAI remains challenging due to an overall lack of transparency by the main actors in the field. Even when the environmental impacts of specific models are mentioned, they are typically restricted to the carbon foot… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/marta-lopez-rauhut/gen-ai-footprint · OpenReview ID:uurX0xsr8G

    最高第 3500:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  42. 42
    Unbiased Stochastic Optimization for Gaussian Processes on Finite Dimensional RKHS

    期刊:Transactions on Machine Learning Research · 摘要:Current methods for stochastic hyperparameter learning in Gaussian Processes (GPs) rely onapproximations, suchascomputingbiasedstochasticgradientsorusinginducingpointsin stochastic variational inference. However, when using such methods, we are not guaranteed to converge to a stationary point of the true marginal likelihood. In this work, we propose algorithms for exact stochastic inference of GPs with kernels that induce a Reproducing Kernel Hilbert Space (RKHS) of moderate finite dimension. Our approach can also be extendedtoinfinitedimensionalRKHSsatthecostofforgoingexactness. Bothforfinit… · 篇幅:Long submission (more than 12 pages of main content) · OpenReview ID:wDCulUZla4

    最高第 1400:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  43. 43
    Transformer–SSM Hybrid Language Models: Systematic Analysis and Design Insights

    期刊:Transactions on Machine Learning Research · 摘要:Recent progress in large language models demonstrates that hybrid architectures--combining self-attention mechanisms with state-space layers--can achieve a compelling balance between modeling quality and computational efficiency, particularly for long-context tasks. While these Transformer–Mamba-2 hybrid models show promising performance, systematic comparisons of hybridization strategies and analyses on the key factors behind their effectiveness have not been clearly shared with the community. In this work, we present a holistic evaluation of hybrid architectures based on inter-layer (sequen… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:x7qyXl8ecT

    最高第 2800:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  44. 44
    Stronger Approximation Guarantees for Non-Monotone $\gamma$-Weakly DR-Submodular Maximization

    期刊:Transactions on Machine Learning Research · 摘要:We study the maximization of nonnegative, non-monotone $\gamma$-weakly diminishing-returns (DR) submodular functions over down-closed convex bodies. The weakly DR model relaxes classical diminishing returns by allowing marginal gains to decay up to a multiplicative factor $\gamma \in (0,1]$, capturing a broad class of objectives that interpolate between monotone and fully non-monotone DR submodularity. Existing methods in this regime achieve guarantees that deteriorate rapidly as $\gamma$ decreases and fail to recover the best known bounds in the fully DR case. We develop a $\gamma$-aware alg… · 篇幅:Long submission (more than 12 pages of main content) · OpenReview ID:yS78Cb1CnX

    最高第 2600:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  45. 45
    EHR2Path: Comprehensive Pathway-Level Modeling of Longitudinal Patient Trajectories from Multimodal Electronic Health Records

    期刊:Transactions on Machine Learning Research · 摘要:Forecasting how a patient’s condition is likely to evolve, including possible deterioration, recovery, treatment needs, and care transitions, could support more proactive and personalized care, but requires modeling heterogeneous and longitudinal electronic health record (EHR) data. Yet, existing approaches typically focus on isolated prediction tasks, narrow feature spaces, or short context windows, limiting their ability to model full patient pathways. To address this gap, we introduce EHR2Path, a multimodal framework for forecasting and simulating full in-hospital patient pathways from rou… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/ChantalMP/EHR2Path · OpenReview ID:ywa71iOykg

    最高第 700:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时49分
  46. 46
    Wiring the ‘Why’: A Unified Taxonomy and Survey of Abductive Reasoning in LLMs

    期刊:Transactions on Machine Learning Research · 摘要:Despite its foundational role in human discovery and sense-making, abductive reasoning—the inference of the most plausible explanation for an observation—has been relatively underexplored in Large Language Models (LLMs). Although LLMs have advanced rapidly, research on abductive reasoning and its diverse facets has remained disjointed rather than cohesive. To the best of our knowledge, this paper presents the first survey dedicated specifically to abductive reasoning in LLMs, tracing its trajectory from philosophical foundations to contemporary LLM-based approaches. To address the widespread… · 篇幅:Long submission (more than 12 pages of main content) · OpenReview ID:oeVkugH0WB

    最高第 4600:00 达到当日首次采集时已在榜23:01 观测离榜累计约23小时2分
  47. 47
    Provably Safe Generative Sampling with Constricting Barrier Functions

    期刊:Transactions on Machine Learning Research · 摘要:Flow-based generative models, such as diffusion models and flow matching models, have achieved remarkable success in learning complex data distributions. However, a critical gap remains for their deployment in safety-critical domains: the lack of formal guarantees that generated samples will satisfy hard constraints. We propose a safety filtering framework that acts as an online shield for any pre-trained generative model. Our key insight is to cooperate with the generative process rather than override it. We define a constricting safety tube that is relaxed at the initial noise distribution… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/darshangm/constricted-diffusion · OpenReview ID:iZi471b4Pf

    最高第 4700:00 达到当日首次采集时已在榜16:05 观测离榜累计约16小时6分
  48. 48
    Neural Diversity Regularizes Hallucinations in Language Models

    期刊:Transactions on Machine Learning Research · 摘要:Language models continue to hallucinate despite scaling parameters, compute, and data. We propose neural diversity — decorrelated parallel representations — as a provable mechanism to reduce hallucination rates at fixed parameter and data budgets. While existing mitigation strategies largely target accuracy, we reframe it as a second-moment reliability problem governed by representational covariance and provide the first formal tail bounds for hallucination probability in ensembled language models, explaining 94.3% of reliability variation across configurations in our setting (Qwen2.5-0.5B, 2… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/kushalc/nd-lora · OpenReview ID:5l9ZflyApA

    最高第 4800:00 达到当日首次采集时已在榜15:17 观测离榜累计约15小时18分
  49. 49
    Task-Relevant Language-conditioned Segmentation for Robust Generalization in Reinforcement Learning

    期刊:Transactions on Machine Learning Research · 摘要:Humans possess a remarkable ability to filter out irrelevant sensory clutter, extracting only the information needed to anticipate and act within dynamic environments. Prior attempts to mitigate this through augmentation and masking strategies have improved robustness, but remain limited by computational overhead, weak semantic grounding, or instability in actor-critic training. Inspired by how language guides human perception, we introduce Task Relevant Language-conditioned Segmentation (TaLaS), a framework that leverages language-conditioned segmentation to impose semantic structure on visu… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:MRHXB6eooE

    最高第 4900:00 达到当日首次采集时已在榜13:57 观测离榜累计约13小时58分
  50. 50
    A Closed-Form Persistence-Landmark Pipeline for Certified Point-Cloud and Graph Classification

    期刊:Transactions on Machine Learning Research · 摘要:We introduce PLACE (Persistence-Landmark Analytic Classification Engine), a closed-form pipeline for classifying point clouds and graphs through their persistent-homology signatures. Three quantitative guarantees—a margin-based excess-risk rate, a closed-form descriptor-selection rule, and a per-prediction certificate—are derived from training labels alone, with no learned weights or held-out calibration. The embedding sums Mitra–Virk single-point coordinate functions over a sparse landmark grid; the closed-form weight rule $w_k^2 \propto (d_{k+1}^2 - d_k^2)/R_k^2$ maximizes the distortion sl… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/akritihq/place-palace · OpenReview ID:4kZxNlE5Ve

    最高第 111:33 达到11:33 首次观测上榜当日结束时仍在榜累计约12小时16分