全部/知识/实时热榜

OpenReview · 实时热榜

HISTORY2026年8月4日60 不同热搜
08/0309/01 有历史数据
DAILY UNIQUE TOPICS60 个热搜
  1. 01
    Reasoning with Preference Constraints: A Benchmark for Language Models in Many-to-One Matching Markets

    期刊:Transactions on Machine Learning Research · 摘要:Recent advances in reasoning with large language models (LLMs) have demonstrated strong performance on complex mathematical tasks. Techniques such as Chain-of-Thought and In-Context Learning have further enhanced this capability, making LLMs both powerful and accessible tools for a wide range of users, including non-experts. However, their application to problems arising in operations research, particularly those at the intersection of combinatorial optimization and game theory that require domain expertise, remains underexplored. To address this gap, we introduce a benchmark of 369 instances… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/maryloufauchard/CAP_Benchmark · OpenReview ID:2dpt2Ughzt

    最高第 600:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  2. 02
    On the Convergence Analysis of Muon

    期刊:Transactions on Machine Learning Research · 摘要:The majority of parameters in neural networks are naturally represented as matrices. However, most commonly used optimizers treat these matrix parameters as flattened vectors during optimization, potentially overlooking their inherent structural properties. Recently, an optimizer called Muon has been proposed, specifically designed to optimize matrix-structured parameters. Extensive empirical evidence shows that Muon can significantly outperform traditional optimizers when training neural networks. Nonetheless, the theoretical understanding of Muon’s convergence behavior and the reasons behin… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:4nH4CulGaP

    最高第 3700:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  3. 03
    Neural Diversity Regularizes Hallucinations in Language Models

    期刊:Transactions on Machine Learning Research · 摘要:Language models continue to hallucinate despite scaling parameters, compute, and data. We propose neural diversity — decorrelated parallel representations — as a provable mechanism to reduce hallucination rates at fixed parameter and data budgets. While existing mitigation strategies largely target accuracy, we reframe it as a second-moment reliability problem governed by representational covariance and provide the first formal tail bounds for hallucination probability in ensembled language models, explaining 94.3% of reliability variation across configurations in our setting (Qwen2.5-0.5B, 2… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/kushalc/nd-lora · OpenReview ID:5l9ZflyApA

    最高第 1100:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  4. 04
    What Survives Privatization? A Guide to Structure and Utility in Differentially Private Genome-Wide Association Studies

    期刊:Transactions on Machine Learning Research · 摘要:Single nucleotide polymorphisms (SNPs) are among the most common and informative forms of genetic variation in the human genome and constitute the primary data representation used in genome-wide association studies (GWAS). Due to their extreme dimensionality, strong correlation structure, and the presence of both population-level and familial dependencies, SNP datasets exhibit structural properties that fundamentally distinguish them from standard tabular data. At the same time, genomic data is uniquely sensitive; it is immutable, identifying, and shared across relatives, and has been shown t… · 篇幅:Long submission (more than 12 pages of main content) · OpenReview ID:6BWikkmkOH

    最高第 2100:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  5. 05
    FreeEyeglass: Training-free and Target-mask-free Eyeglass Transfer for Facial Videos

    期刊:Transactions on Machine Learning Research · 摘要:The rise of e-commerce and short-video platforms has fueled demand for realistic video-based virtual try-on. Unlike virtual try-on of clothing, which has been actively studied to date, virtual try-on of eyeglasses is uniquely challenging: they align closely with facial structure and strongly affect facial identity, making the faithful preservation of unedited regions especially important. Existing generative editing approaches, such as GAN- and diffusion-based methods, lack reconstruction objectives and often rely on inpainting, which fails to ensure identity consistency. We argue that semant… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:6aFRoQcm3H

    最高第 3900:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  6. 06
    CDG-MAE: Cross-view Masked Modeling using Diffusion Generated Views

    期刊:Transactions on Machine Learning Research · 摘要:Cross-view masked autoencoding has emerged as a powerful pretext task for learning dense correspondences, which are essential for applications such as video label propagation. The cross-view pretext task is modeled with a masked autoencoder, where a masked target view is reconstructed from an anchor view. However, acquiring effective training data remains a challenge - collecting diverse video datasets is costly, while simple image crops lack the necessary pose variations, underperforming video-based methods. This paper introduces CDG-MAE, a novel MAE-based self-supervised method that uses di… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/cvlab-stonybrook/CDG-MAE · OpenReview ID:7XIymKIA0v

    最高第 1800:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  7. 07
    Dynamic Reward Incentives for Emergent Cooperation under Changing Rewards

    期刊:Transactions on Machine Learning Research · 摘要:Peer incentivization (PI) is a popular multi-agent reinforcement learning approach where all agents can reward or penalize each other to achieve cooperation in social dilemmas. Despite their potential for scalable cooperation, current PI methods heavily depend on fixed incentive values that need to be appropriately chosen with respect to the environmental rewards and thus are highly sensitive to their changes. Therefore, they fail to maintain cooperation under changing rewards in the environment, e.g., caused by modified specifications, varying supply and demand, or sensory flaws — even when… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/philippaltmann/DRIVE · OpenReview ID:9Ltu1HV2YI

    最高第 400:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  8. 08
    NeMoS: Nearest Neighbors Bandit meets Active Learning for Online Model Selection

    期刊:Transactions on Machine Learning Research · 摘要:The proliferation of open-platform text-to-image generative models has made prompt-wise model selection critical to maximize generation quality and semantic alignment. However, current strategies, such as contextual bandits, often converge slowly and fail to exploit the semantic relationships across prompts. To bridge this gap, we propose NeMoS, a non-parametric bandit framework that couples nearest neighbor reward estimation with a budget-constrained active learning strategy. Specifically, our approach operates in the prompt embedding space and estimates the reward of incoming prompts based… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/julesdamidaux/nemos-tmlr · OpenReview ID:CSjewjplO1

    最高第 800:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  9. 09
    PAC-Bayesian Meta-Learning for Few-Shot Identification of Linear Dynamical Systems

    期刊:Transactions on Machine Learning Research · 摘要:Identifying linear time-invariant (LTI) dynamical systems from data is especially challenging when trajectories are short, noisy, or high-dimensional. Traditional system identification methods typically treat each system in isolation and therefore fail to exploit shared structure across related systems. We propose a PAC-Bayesian meta-learning framework for few-shot LTI system identification (PBML-LTI), which learns a transferable prior over task-specific dynamics while preserving task-level heterogeneity. Each task corresponds to an unknown LTI system, and a meta-learner uses a collection of… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/chenfeng-huang/PBML-LTI-TMLR-2026 · OpenReview ID:CiGFpSLzFv

    最高第 500:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  10. 10
    On Almost Surely Safe Alignment of Large Language Models at Inference Time

    期刊:Transactions on Machine Learning Research · 摘要:We introduce a novel inference-time alignment approach for LLMs that aims to generate safe responses almost surely, i.e., with probability approaching one w.r.t. a given cost model. Our approach models the generation of safe responses as a constrained Markov Decision Process (MDP) within the LLM's latent space. We augment a safety state that tracks the evolution of safety constraints and dynamically penalize unsafe generations to ensure the generation of safe responses. Consequently, we demonstrate formal safety guarantees w.r.t. the given cost model upon solving the MDP in the latent space w… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/rsshyam/inf-guard · OpenReview ID:FlnokjaSEu

    最高第 2500:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  11. 11
    Watermarking Language Models with Error Correcting Codes

    期刊:Transactions on Machine Learning Research · 摘要:Recent progress in large language models enables the creation of realistic machine-generated content. Watermarking is a promising approach to distinguish machine-generated text from human text, embedding statistical signals in the output that are ideally undetectable to humans. We propose a watermarking framework that encodes such signals through an error correcting code. Our method, termed robust binary code (RBC) watermark, introduces no noticeable degradation in quality. We evaluate our watermark on base and instruction fine-tuned models and find that our watermark is robust to edits, dele… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/patrickrchao/watermarking-llms · OpenReview ID:H6oBZxNQk2

    最高第 2300:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  12. 12
    Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing

    期刊:Transactions on Machine Learning Research · 摘要:Quadratic self-attention remains a central obstacle to scaling sequence length, but many sparse attention methods lose quality relative to dense attention under comparable compute budgets. This paper studies whether learned, content-based tokezzn selection can make sparse attention competitive with dense attention as a method for training transformer language models. We present Mixture of Sparse Attention (MoSA), an attention mechanism inspired by Mixture of Experts with expert-choice routing, where each attention head selects its own subset of tokens and computes attention only within that s… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:HUpBs4TZkS

    最高第 3800:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  13. 13
    Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning

    期刊:Transactions on Machine Learning Research · 摘要:Learning general-purpose reasoning capabilities has long been a challenging problem in AI. Recent research in LLMs, such as DeepSeek-R1, has shown that reinforcement learning techniques like GRPO enable pre-trained LLMs to develop reasoning capabilities using simple question-answer pairs. In this paper, we aim to train visual language models (VLMs) to perform reasoning on image data through reinforcement learning and visual question-answer pairs, without explicitly using any chain-of-thought (CoT) supervision. Our key finding indicates that simply applying GRPO to a VLM---by prompting the mod… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/maifoundations/Visionary-R1 · OpenReview ID:JWkZXBgh5a

    最高第 700:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  14. 14
    ViP$^2$-CLIP: Visual-Perception Prompting with Unified Alignment for Zero-Shot Anomaly Detection

    期刊:Transactions on Machine Learning Research · 摘要:Zero-Shot Anomaly Detection (ZSAD) aims to detect anomalies in a target dataset without any training samples, leveraging models trained on auxiliary data. While CLIP offers strong cross-modal representations for ZSAD, its pretraining objective inherently emphasizes global foreground semantics over fine-grained local defects. Consequently, its anomaly localization remains highly sensitive to prompt wording, limiting the effectiveness of existing methods that rely on explicit category labels. To overcome this limitation, we introduce ViP$^{2}$-CLIP, a lightweight CLIP-based ZSAD framework featu… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:KCRRuiQSIm

    最高第 1500:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  15. 15
    Unified Sample Difficulty Estimation in Pathology Foundation Models

    期刊:Transactions on Machine Learning Research · 摘要:The fast scaling speed of histopathology datasets allows researchers to train various foundation models for disease-centered research with applications in classifying disease-state information and predicting gene expression levels. However, it has been shown that current models tend to be overconfident and make classification at a low-calibration level. This case is underexplored for regression-type tasks such as gene expression prediction as well, which could seriously affect the diagnosis and treatment based on the developed models. To resolve this critical issue, we propose a \underline{u}… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:LLlOJs4o2N

    最高第 2000:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  16. 16
    Task-Relevant Language-conditioned Segmentation for Robust Generalization in Reinforcement Learning

    期刊:Transactions on Machine Learning Research · 摘要:Humans possess a remarkable ability to filter out irrelevant sensory clutter, extracting only the information needed to anticipate and act within dynamic environments. Prior attempts to mitigate this through augmentation and masking strategies have improved robustness, but remain limited by computational overhead, weak semantic grounding, or instability in actor-critic training. Inspired by how language guides human perception, we introduce Task Relevant Language-conditioned Segmentation (TaLaS), a framework that leverages language-conditioned segmentation to impose semantic structure on visu… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:MRHXB6eooE

    最高第 1200:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  17. 17
    Cross-Fitted Clipped Covariance Estimation with a Data-Driven Tail-Energy Criterion

    期刊:Transactions on Machine Learning Research · 摘要:Heavy-tailed data make covariance estimation sensitive to the clipping level: stronger clipping reduces variance but increases bias. We study how to choose this clipping level from the data within a radial clipped covariance family. We propose the quantile tail-energy surrogate (QTES), a fully data-driven rule that combines a cross-fitted variance certificate with a held-out estimate of the tail energy removed by clipping. QTES requires no distributional prior parameters. For Euclidean clipping, the operator-norm bias is bounded by this scalar tail-energy quantity. Under a finite $L_4$ moment… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:MyNXLdRFJ3

    最高第 1400:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  18. 18
    RIGID: A Training-Free and Generator-Agnostic Framework for Robust AI-Generated Image Detection

    期刊:Transactions on Machine Learning Research · 摘要:The rapid advances in generative AI models have empowered the creation of highly realistic images with arbitrary content, raising concerns about potential misuse and harm, such as Deepfakes. Current research focuses on training detectors using large datasets of generated images. However, these training-based solutions are often computationally expensive and show limited generalization to unseen generated images. In this paper, we propose a training-free method to distinguish between real and AI-generated images. We first observe that real images are more robust to tiny noise perturbations tha… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/IBM/RIGID · OpenReview ID:NBkBI2Zjlm

    最高第 1600:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  19. 19
    GEO-Bench-2: From Performance to Capability, Rethinking Evaluation in Geospatial AI

    期刊:Transactions on Machine Learning Research · 摘要:Geospatial Foundation Models (GeoFMs) are transforming Earth Observation (EO), but evaluation lacks standardized protocols. GEO-Bench-2 addresses this with a com- prehensive framework spanning classification, segmentation, regression, object detection, and instance segmentation across 19 permissively-licensed datasets. We introduce capabil- ity groups to rank models on datasets that share common characteristics (e.g., resolution, spectral bands, temporality), enabling users to identify which models excel in each capa- bility and to determine where future work should focus. To support both fai… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:NPf175jnP1

    最高第 3500:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  20. 20
    FUND: Density Flow for Sampling Unnormalised Distributions

    期刊:Transactions on Machine Learning Research · 摘要:Efficient sampling from Boltzmann distributions is central to modelling complex physical systems. Markov Chain Monte Carlo (MCMC) methods suffer from critical slowing down, high autocorrelation, and poor mode-mixing, limiting their scalability. Recent advances, like Boltzmann Generators, offer a promising alternative but remain constrained by costly MCMC-based training, inefficient sampling, and poor ergodicity. We introduce an algorithm for learning Boltzmann distributions that does not require any true samples for training. Our approach draws inspiration from flow matching but departs funda… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/madhavlab/2025_fund · OpenReview ID:O05dDDVcyZ

    最高第 2900:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  21. 21
    On-the-go Forgetting without Explicit Unlearning via ERASE

    期刊:Transactions on Machine Learning Research · 摘要:Existing unlearning approaches typically rely on post hoc weight adaptation or distillation, leading to duplicated memory costs, degraded generalization, and limited scalability. In this work, we introduce ERASE, Erasure via Reconstructive Adversarial Signal Editing, a framework for on-the-go forgetting that suppresses the observable influence of private data without modifying model weights. ERASE leverages structured, class-conditioned input perturbations to induce selective forgetting during inference, eliminating the need for retraining, fine-tuning, or model copies. We rigorously characte… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:PIXVov5LQq

    最高第 2200:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  22. 22
    Revisiting Neighbourhoods in Mean Field Reinforcement Learning

    期刊:Transactions on Machine Learning Research · 摘要:Many multi-agent reinforcement learning (MARL) algorithms do not scale well as the number of agents increases due to an exponential time and space complexity dependency on the number of agents in the environment. Mean field theory has been used to address this problem by approximating the effect of neighbourhoods of agents by a single representative agent. While this approximation allows MARL algorithms to scale to environments with many agents, approaches typically assumed that agents 1) inside a neighbourhood are homogeneous, and 2) outside a neighbourhood have no influence (and can therefo… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/Sriram94/MFA · OpenReview ID:PQ5R7K0WDc

    最高第 3600:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  23. 23
    Automata Learning from Recurrent Networks: A Critical Synthesis for Verification, Testing, and Interpretability

    期刊:Transactions on Machine Learning Research · 摘要:Recurrent Neural Networks (\RNNs) have demonstrated their effectiveness in modeling sequential data and are a key building block of modern deep learning architectures. In this review paper, we study recurrent networks through the lens of automata theory. Given an \RNN, automata learning seeks to model its behavior with an automaton, which enables better interpretability and eases our understanding of its working mechanisms. We begin by examining the theoretical foundations of this approach, demonstrating how it can be applied to learn automata from various types of recurrent architectures, in… · 篇幅:Long submission (more than 12 pages of main content) · OpenReview ID:R52ETbUBVo

    最高第 3200:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  24. 24
    Revisiting Learning-based Video Motion Magnification for Real-time Processing

    期刊:Transactions on Machine Learning Research · 摘要:Video motion magnification is a technique to capture and amplify subtle motion in a video that is invisible to the naked eye. The deep learning-based prior work successfully models outstanding quality better than conventional signal processing-based ones. However, it still lags behind real-time performance, which prevents it from being extended to various online systems. In this paper, we revisit the first learning-based model and present experimental analyses, in particular on the identification of redundant components, the insertion of spatial bottlenecks, and the trade-off relationship bet… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/kaist-ami/Fast-MM · OpenReview ID:TAmmPuExE1

    最高第 2700:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  25. 25
    Benford’s Law as a Distributional Prior for Post-Training Quantization of Large Language Models

    期刊:Transactions on Machine Learning Research · 摘要:Post-training quantization (PTQ) is a practical way to reduce the memory footprint of large language models, but low-bit quantization is sensitive to mismatches between the quantization codebook and the empirical weight/activation distributions. We revisit Benford-like leading-digit statistics as a lightweight diagnostic of scale-broad behavior in transformer tensors. Across several model families, we observe a consistent functional dichotomy: transformational nn.Linear weights tend to be Benford-like, whereas LayerNorm and embedding parameters systematically deviate. Motivated by this observ… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/ufopcsilab/benford-quant · OpenReview ID:YiLcQY4Nje

    最高第 1900:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  26. 26
    Trade-offs in Ensembling, Merging and Routing Among Parameter-Efficient Experts

    期刊:Transactions on Machine Learning Research · 摘要:While large language models (LLMs) fine-tuned with lightweight adapters achieve strong performance across diverse tasks, their performance on individual tasks depends on the fine-tuning strategy. Fusing independently trained models with different strengths has shown promise for multi-task learning through three main strategies: ensembling, which combines outputs from independent models; merging, which fuses model weights via parameter averaging; and routing, which integrates models in an input-dependent fashion. However, many design decisions in these approaches remain understudied, and the r… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:bnRCvRtZv5

    最高第 2800:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  27. 27
    Autoregressive Image Generation with Frequency Progression

    期刊:Transactions on Machine Learning Research · 摘要:Autoregressive (AR) models for image generation typically adopt a two-stage paradigm of vector quantization and raster-scan ``next-token prediction", inspired by its great success in language modeling. However, due to the huge modality gap, image autoregressive models may require a systematic reevaluation from two perspectives: tokenizer format and regression direction. In this paper, we introduce the frequency progressive autoregressive (\textbf{FAR}) paradigm and instantiate FAR with the continuous tokenizer. Specifically, we identify spectral dependency as the desirable regression directio… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:cEfd15ouQ1

    最高第 4000:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  28. 28
    Learning 3D Hypersonic Flow with Physics-Enhanced Neural Fields: A Case Study on the Orion Reentry Capsule

    期刊:Transactions on Machine Learning Research · 摘要:We develop a 3D aerothermodynamic surrogate for the Orion reentry capsule at hypersonic speeds, a timely case study given its role in upcoming lunar missions. The large computational meshes required for these scenarios make traditional computational fluid dynamics impractical for full-mission performance prediction and control. In this work, we propose physics-enhanced 3D neural fields for predicting steady hypersonic flow around aerodynamic bodies. The model maps spatial coordinates and angle of attack to pressure, temperature, and velocity components. We enhance the base model with Fourier… · 篇幅:Long submission (more than 12 pages of main content) · OpenReview ID:ce2X1X3l0Y

    最高第 3400:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  29. 29
    Beyond Naïve Prompting: Strategies for Improved Context-aided Forecasting with LLMs

    期刊:Transactions on Machine Learning Research · 摘要:Real-world forecasting requires models to integrate not only historical data but also relevant contextual information provided in textual form. While large language models (LLMs) show promise for context-aided forecasting, critical challenges remain: we lack diagnostic tools to understand failure modes, performance remains far below their potential, and high computational costs limit practical deployment. We introduce a unified framework of four strategies that address these limitations along three orthogonal dimensions: model diagnostics, accuracy, and efficiency. Through extensive evaluatio… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/ashok-arjun/beyond-naive-prompting · OpenReview ID:dkjHHFJkVI

    最高第 3100:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  30. 30
    Quantification and Control of LSTM Resilience Based on Stability Theory

    期刊:Transactions on Machine Learning Research · 摘要:This paper proposes a novel theoretical framework for guaranteeing and evaluating the resilience of long short-term memory (LSTM) networks in control systems. We introduce *recovery time* as a new metric of resilience in order to quantify the time required for an LSTM to return to its normal state after anomalous inputs. By mathematically refining incremental input-to-state stability ($\delta$ISS) theory for LSTM, we derive a practical data-independent upper bound on recovery time. This upper bound gives us resilience-aware training. Experimental validation on simple models demonstrates the e… · 篇幅:Regular submission (no more than 12 pages of main content) · OpenReview ID:hFmlMUNEsR

    最高第 3300:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  31. 31
    Provably Safe Generative Sampling with Constricting Barrier Functions

    期刊:Transactions on Machine Learning Research · 摘要:Flow-based generative models, such as diffusion models and flow matching models, have achieved remarkable success in learning complex data distributions. However, a critical gap remains for their deployment in safety-critical domains: the lack of formal guarantees that generated samples will satisfy hard constraints. We propose a safety filtering framework that acts as an online shield for any pre-trained generative model. Our key insight is to cooperate with the generative process rather than override it. We define a constricting safety tube that is relaxed at the initial noise distribution… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/darshangm/constricted-diffusion · OpenReview ID:iZi471b4Pf

    最高第 1000:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  32. 32
    Federated Class-Incremental Learning with Hierarchical Generative Prototypes

    期刊:Transactions on Machine Learning Research · 摘要:Federated Learning (FL) aims at unburdening the training of deep models by distributing computation across multiple devices (clients) while safeguarding data privacy. On top of that, Federated Continual Learning (FCL) also accounts for data distribution evolving over time, mirroring the dynamic nature of real-world environments. While previous studies have identified Catastrophic Forgetting and Client Drift as primary causes of performance degradation in FCL, we shed light on the importance of Incremental Bias and Federated Bias, which cause models to prioritize classes that are recently intr… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/aimagelab/fed-mammoth · OpenReview ID:k2TT42Ei8W

    最高第 300:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  33. 33
    DIMENSION DOMAIN CO-DECOMPOSITION: SOLVING PDES WITH INTERPRETABILITY

    期刊:Transactions on Machine Learning Research · 摘要:Physics-informed neural networks (PINNs) have demonstrated effectiveness in solving partial differential equations (PDEs), yet they often struggle in high-dimensional regimes and lack interpretable representations and in scenarios involving sharp solution structures. Moreover, existing approaches typically rely on manually specified domain partitions. We propose a unified Dimension–Domain Co-Decomposition (3D) framework that jointly integrates dimension-wise decomposition with mixture-of-experts (MoE)–based domain decomposition. At the dimension level, we introduce an interpretable decomposit… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/TIML-Group/3DPINN · OpenReview ID:kuzkynVyRq

    最高第 200:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  34. 34
    ARC-Encoder: learning compressed text representations for large language models

    期刊:Transactions on Machine Learning Research · 摘要:Recent techniques such as retrieval-augmented generation or chain-of-thought reasoning have led to longer contexts and increased inference costs. Context compression techniques can reduce these costs, but the most effective approaches require fine-tuning the target model or even modifying its architecture. This can degrade its general abilities when not used for this specific purpose. Here we explore an alternative approach: an encoder that compresses the context into continuous representations which replace token embeddings in decoder LLMs. First, we perform a study of training strategies an… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/kyutai-labs/ARC-Encoder · OpenReview ID:lU1P9dsqfn

    最高第 3000:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  35. 35
    ANU-RL: A New Perspective on Weakly-Supervised Representation Learning for Visual Place Recognition

    期刊:Transactions on Machine Learning Research · 摘要:Representation Learning (RL) is fundamental for image matching, retrieval, classification, and other applications, enabling task-specific feature learning. RL algorithms aim to learn compact embeddings that preserve the neighbourhood structure of the input data. A general approach to this is contrastive learning, which pulls similar images (positives) closer together and pushes dissimilar images (negatives) farther apart in the embedding space. In Visual Place Recognition (VPR), positive images of a query share specific geographical and visual attributes with the query and can form a cluster.… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/Anuradha-Uggi/ANU-RL · OpenReview ID:mXE4OP55il

    最高第 100:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  36. 36
    Verify What Matters: Budgeted Verification for Tool-Using Agents under Counterfactual Downstream Harm

    期刊:Transactions on Machine Learning Research · 摘要:Tool-using agents make intermediate decisions that alter persistent state, shape later observations, and create failures that are not equally easy to recover from. When verification is costly, the central question is not whether checking helps in general, but which decisions are worth checking. Policies driven only by local uncertainty capture whether a step may be wrong, but not how much that error would matter if left uncorrected. We formulate budgeted verification for tool-using agents as an intervention-allocation problem in which the value of checking a step depends on verifier efficacy,… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/tang03130313/verify-what-matters · OpenReview ID:nv1jzr0FaZ

    最高第 1300:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  37. 37
    Wiring the ‘Why’: A Unified Taxonomy and Survey of Abductive Reasoning in LLMs

    期刊:Transactions on Machine Learning Research · 摘要:Despite its foundational role in human discovery and sense-making, abductive reasoning—the inference of the most plausible explanation for an observation—has been relatively underexplored in Large Language Models (LLMs). Although LLMs have advanced rapidly, research on abductive reasoning and its diverse facets has remained disjointed rather than cohesive. To the best of our knowledge, this paper presents the first survey dedicated specifically to abductive reasoning in LLMs, tracing its trajectory from philosophical foundations to contemporary LLM-based approaches. To address the widespread… · 篇幅:Long submission (more than 12 pages of main content) · OpenReview ID:oeVkugH0WB

    最高第 900:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  38. 38
    Accelerating SGDM via Learning Rate and Batch Size Schedules: A Lyapunov-Based Analysis

    期刊:Transactions on Machine Learning Research · 摘要:We analyze the convergence behavior of stochastic gradient descent with momentum (SGDM) under dynamic learning-rate and batch-size schedules by introducing a novel and simpler Lyapunov function. We extend the existing theoretical framework to cover three practical scheduling strategies commonly used in deep learning: a constant batch size with a decaying learning rate, an increasing batch size with a decaying learning rate, and an increasing batch size with an increasing learning rate. Our results reveal a clear hierarchy in convergence: a constant batch size does not guarantee convergence of… · 篇幅:Long submission (more than 12 pages of main content) · OpenReview ID:s6DTv7Sorj

    最高第 2600:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  39. 39
    FieldFormer: Locality-Aware Transformers for Spatio-Temporal Modeling on Sparse Sensor Networks

    期刊:Transactions on Machine Learning Research · 摘要:Spatio-temporal sensor data in real-world systems is often sparse, noisy, and irregular, making it difficult to infer global structure from limited observations. Under extreme sparsity, we run into the limits of identifiability of latent system states, making latent field reconstruction fundamentally underconstrained. In such scenarios, multiple physically plausible fields may remain consistent with the same observations, requiring reconstruction models to rely heavily on inductive biases regarding locality, transport structure, and spatial regularity. Under such sparsity regimes, reliable re… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/ankitbha/fieldformer · OpenReview ID:we4FYGOE2y

    最高第 1700:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  40. 40
    Random features for Grassmannian kernel approximation with bounded rank-one projections

    期刊:Transactions on Machine Learning Research · 摘要:We propose a family of random feature maps for scalable kernel machines defined over low-dimensional subspaces in high dimensions, \ie over the Grassmannian manifold. This is typically useful in a machine learning context when data classes or clusters are well represented by the span of a few data points. Classical Grassmannian kernels such as the \emph{projection} or \emph{Binet–Cauchy} kernels require constructing full Gram matrices for practical applications, leading to prohibitive computational and memory costs for large subspace datasets in high dimensions. We address this limitation by… · 篇幅:Long submission (more than 12 pages of main content) · OpenReview ID:wq18dZJ2pA

    最高第 2400:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时50分
  41. 41
    Optimized Graph Structures for Calibrating Graph Neural Networks with Out-of-Distribution Nodes

    期刊:Transactions on Machine Learning Research · 摘要:edges between ID and OOD nodes can substantially improve calibration. Identifying these edges and assigning appropriate weights is challenging because the identities of OOD nodes are unknown. To address this challenge, we propose Graph Calibration via Structure Optimization (GCSO), a novel framework for calibrating GNNs in the presence of OOD nodes. GCSO introduces an iterative edge-sampling mechanism to capture graph topological information and formulates adaptive structure optimization as a Markov Decision Process (MDP). An actor-critic policy then dynamically adjusts edge weights based on… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/DamoSWL/GCSO · OpenReview ID:Y1W3Z3Z6i8

    最高第 4100:00 达到当日首次采集时已在榜23:33 观测离榜累计约23小时34分
  42. 42
    Sharpness-Aware Minimization Driven by Local-Integrability Flatness

    期刊:Transactions on Machine Learning Research · 摘要:Sharpness-Aware Minimization (SAM) improves generalization by optimizing for worst-case loss under parameter perturbations, but its max-based objective can be overly conservative, noise-sensitive, and reliant on smoothness assumptions that often fail in modern nonsmooth networks. We propose Lebesgue Sharpness-Aware Minimization (LSAM), a measure-theoretic alternative grounded in the Lebesgue Differentiation Theorem and local Sobolev regularity. Instead of minimizing the worst-case loss, LSAM minimizes the local average loss in a neighborhood of the parameters. This average-case notion of flat… · 篇幅:Long submission (more than 12 pages of main content) · OpenReview ID:29Zg9k5NCo

    最高第 4200:00 达到当日首次采集时已在榜23:17 观测离榜累计约23小时18分
  43. 43
    VideoScore2: Think Before You Score In Generated Video Evaluation

    期刊:Transactions on Machine Learning Research · 摘要:Recent advances in text-to-video generation have produced increasingly realistic and diverse content, yet evaluating such videos remains a fundamental challenge due to their multi-faceted nature encompassing visual quality, semantic alignment, and physical consistency. Existing evaluators and reward models are limited to single opaque scores, lack interpretability, or provide only coarse analysis, making them insufficient for capturing the comprehensive nature of video quality assessment. We present VideoScore2, a multi-dimensional, interpretable, and human-aligned framework that explicitly e… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/TIGER-AI-Lab/VideoScore2 · OpenReview ID:MpkVh4jH44

    最高第 4300:00 达到当日首次采集时已在榜23:01 观测离榜累计约23小时2分
  44. 44
    Probing and Controlling Self-Reflection in Language Models

    期刊:Transactions on Machine Learning Research · 摘要:Self-reflection, the ability of a large language model (LLM) to revisit, evaluate, and revise its own reasoning, has recently emerged as a powerful behavior enabled by reinforcement learning with verifiable rewards (RLVR). While self-reflection correlates with improved reasoning accuracy, its origin and underlying mechanisms remain poorly understood. In this work, we first show that self-reflection is not exclusive to RLVR fine-tuned models: it already emerges, albeit rarely, in pretrained models. To probe this latent ability, we introduce Reflection-Inducing Probing, a method that injects re… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/xzAscC/ProbingReflection · OpenReview ID:AwVIfBZwy0

    最高第 4400:00 达到当日首次采集时已在榜22:45 观测离榜累计约22小时46分
  45. 45
    On the Fundamental Limits of LLMs at Scale

    期刊:Transactions on Machine Learning Research · 摘要:Large Language Models (LLMs) have benefited enormously from scaling, yet these gains are bounded by five fundamental limitations: (1) hallucination, (2) context compression, (3) reasoning degradation, (4) retrieval fragility, and (5) multimodal misalignment. While existing surveys describe these phenomena empirically, they lack a rigorous theoretical synthesis connecting them to the foundational limits of computation, information, and learning. This work closes that gap by presenting a unified, proof-informed framework that formalizes the innate theoretical ceilings of LLM scaling. First, com… · 篇幅:Long submission (more than 12 pages of main content) · OpenReview ID:BIRDGVrom8

    最高第 102:30 达到02:30 首次观测上榜当日结束时仍在榜累计约21小时19分
  46. 46
    Inference-Time Computations for LLM Reasoning and Planning: A Benchmark and Insights

    期刊:Transactions on Machine Learning Research · 摘要:We examine the reasoning and planning capabilities of large language models (LLMs) in solving complex tasks. Recent advances in inference-time techniques demonstrate the potential to enhance LLM reasoning without additional training by exploring intermediate steps during inference. Here, we explore how scaling inference-time techniques can improve reasoning and planning, focusing on understanding the tradeoff between computational cost and performance. To this end, we construct a comprehensive benchmark, known as *Sys2Bench*, and perform extensive experiments evaluating existing inference-tim… · 篇幅:Regular submission (no more than 12 pages of main content) · 代码:https://github.com/divelab/sys2bench · OpenReview ID:budZJyCK8G

    最高第 102:46 达到02:46 首次观测上榜当日结束时仍在榜累计约21小时3分
  47. 47
    Conditional Independence Tests for Constraint-Based Causal Discovery: A Survey

    期刊:Transactions on Machine Learning Research · 摘要:Conditional Independence (CI) tests are the statistical engine of constraint-based causal discovery: in algorithms such as PC (Peter-Clark) and FCI (Fast Causal Inference), skeleton pruning and key orientations follow directly from CI decisions. This survey reviews CI testing with emphasis on assumptions, robustness, and scalability in high-dimensional and heterogeneous settings common in biomedical domains. The survey organizes widely used CI methods into six families: partial-correlation, contingency-table, regression, nearest-neighbor, kernel, and machine-learning-based. Special emphasis i… · 篇幅:Long submission (more than 12 pages of main content) · OpenReview ID:3jzafJK8Tz

    最高第 4500:00 达到当日首次采集时已在榜19:01 观测离榜累计约19小时2分
  48. 48
    Environmental Footprint of GenAI Research: Insights from the Moshi Foundation Model

    期刊:Transactions on Machine Learning Research · 摘要:New multi-modal large language models (MLLMs) are continuously being trained and deployed, following rapid development cycles. This generative AI frenzy is driving steady increases in energy consumption, greenhouse gas emissions, and a plethora of other environmental impacts linked to datacenter construction and hardware manufacturing. Mitigating the environmental consequences of GenAI remains challenging due to an overall lack of transparency by the main actors in the field. Even when the environmental impacts of specific models are mentioned, they are typically restricted to the carbon foot… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/marta-lopez-rauhut/gen-ai-footprint · OpenReview ID:uurX0xsr8G

    最高第 106:14 达到06:14 首次观测上榜当日结束时仍在榜累计约17小时35分
  49. 49
    FedIndex: Federated Domain Adaptation with Continuous Domain Indices

    期刊:Transactions on Machine Learning Research · 摘要:Federated domain adaptation incorporates source clients’ knowledge to improve the model performance on the target client under the coordination of the server, mitigating the impact of data insufficiency and domain shift. Existing federated domain adaptation (FDA) methods focus on domain adaptation with categorical domain indices (e.g., “source” and “target”), while many real-world tasks involve domains with continuous domain indices. For instance, hospitals need to adapt disease analysis and prediction across patients via age, a continuous domain index in medical applications capturing the un… · 篇幅:Long submission (more than 12 pages of main content) · 代码:https://github.com/IntelliSys-Lab/FedIndex · OpenReview ID:fnbGFH0330

    最高第 4600:00 达到当日首次采集时已在榜14:13 观测离榜累计约14小时14分
  50. 50
    How Much Information Fits in a Vector?

    期刊:Transactions on Machine Learning Research · 摘要:Recent work in neural network interpretability has suggested that hidden activations of some deep models can be viewed as linear projections of much higher-dimensional vectors of sparse latent ``features.'' In general, this kind of representation is known as a superposition code. This work presents an information-theoretic account of superposition codes in a setting applicable to interpretability. We show that when the number $k$ of active features is very small compared to the number $N$ of total features, simple inference methods currently used by sparse autoencoders can reliably decode a $… · 篇幅:Long submission (more than 12 pages of main content) · OpenReview ID:Nby4pCPIZI

    最高第 111:34 达到11:34 首次观测上榜当日结束时仍在榜累计约12小时15分