东方财富东方财富华尔街见闻华尔街见闻同花顺同花顺微信读书微信读书LichessLichessV2EXV2EX微博微博今日头条今日头条百度百度快手快手哔哩哔哩哔哩哔哩知乎知乎腾讯新闻腾讯新闻网易新闻网易新闻澎湃新闻澎湃新闻新浪新闻新浪新闻新浪网新浪网豆瓣电影豆瓣电影GitHubGitHubCSDNCSDNIT之家IT之家36氪36氪宝庆银楼宝庆银楼中国黄金中国黄金周生生周生生英皇珠宝英皇珠宝老凤祥(广东)老凤祥(广东)六福珠宝六福珠宝周大福周大福周六福周六福AcFunAcFunAdafruit BlogAdafruit BlogAIbaseAIbaseAnt Bailing Developer BlogAnt Bailing Developer BlogAnthropicAnthropic小众软件小众软件AppleAppleApple PodcastsApple PodcastsAppleInsiderAppleInsiderApp StoreApp StoreArs TechnicaArs Technica汽车媒体榜汽车媒体榜AxiosAxiosBerkeley AI ResearchBerkeley AI Research半月谈半月谈BarchartBarchartBBC NewsBBC NewsBBC SportBBC SportBD Tech TalksBD Tech TalksBerkeley RDIBerkeley RDIBig ThinkBig Think哔哩哔哩热搜哔哩哔哩热搜哔哩哔哩热门视频哔哩哔哩热门视频哔哩哔哩哔哩哔哩BinanceBinance新京报新京报BloombergBloombergBlueskyBlueskyBoston Dynamics BlogBoston Dynamics BlogBusiness InsiderBusiness Insider商业科技媒体商业科技媒体ByteDance Seed ResearchByteDance Seed Research财新网与财经网财新网与财经网参考消息参考消息央视网央视网中国新闻网中国新闻网虫部落虫部落Claude BlogClaude BlogClaude Code ReleasesClaude Code ReleasesCloudflare BlogCloudflare Blog财联社财联社CMU Machine Learning BlogCMU Machine Learning Blog国内科技媒体国内科技媒体博客园 / 开源中国博客园 / 开源中国CoinGeckoCoinGeckoCointelegraphCointelegraph酷安酷安crates.iocrates.ioCrowd SupplyCrowd SupplyCSDNCSDNCSS-TricksCSS-Tricks51CTO51CTOCult of MacCult of MacCursor BlogCursor Blogdaily.dev · Populardaily.dev · PopularDaring FireballDaring FireballDario AmodeiDario AmodeiDeepLearning.AI · The BatchDeepLearning.AI · The BatchGoogle DeepMind BlogGoogle DeepMind BlogDeepSeek GitHubDeepSeek GitHubDefiLlamaDefiLlamaDescript BlogDescript Blog设计社区作品榜设计社区作品榜DEV.toDEV.to数字尾巴数字尾巴Docker HubDocker Hub懂球帝懂球帝豆瓣读书豆瓣读书豆瓣读书豆瓣读书豆瓣豆瓣豆瓣讨论豆瓣讨论Dwarkesh PatelDwarkesh Patel东方财富网东方财富网The EconomistThe EconomistEleutherAI BlogEleutherAI Blog英文科学、体育与综合媒体 feed英文科学、体育与综合媒体 feed英文科技媒体分类 feed英文科技媒体分类 feed英文科技评测与行业媒体 feed英文科技评测与行业媒体 feedEngadgetEngadget中央媒体电子报中央媒体电子报交易所、监管与证券报交易所、监管与证券报Financial TimesFinancial TimesFlathubFlathubfreeCodeCampfreeCodeCampFrontiersFrontiers游戏与数码论坛游戏与数码论坛游戏媒体与平台游戏媒体与平台Gary MarcusGary Marcus极客公园极客公园原神原神果核剥壳果核剥壳Gitee 与 GitLabGitee 与 GitLabGitHubGitHubGitHub 周刊仓库GitHub 周刊仓库GoogleGoogleGoogle TrendsGoogle Trends果壳果壳HackadayHackadayHacker NewsHacker NewsHelloGitHubHelloGitHub历史上的今天历史上的今天HomebrewHomebrew崩坏3崩坏3Hugging FaceHugging Face活动行活动行虎扑虎扑虎扑社区虎扑社区虎嗅虎嗅IEEE SpectrumIEEE Spectrum爱范儿爱范儿凤凰网 · 热点资讯凤凰网 · 热点资讯IMDbIMDbinclusionAIinclusionAIIndie HackersIndie HackersInfoQInfoQ南方周末南方周末InterconnectsInterconnects投资社区与研究报告投资社区与研究报告爱奇艺热播榜爱奇艺热播榜爱奇艺风云榜爱奇艺风云榜IT之家与快科技IT之家与快科技IT之家「喜加一」IT之家「喜加一」简书简书稀土掘金稀土掘金掘金 / InfoQ掘金 / InfoQAndrej KarpathyAndrej KarpathyMoonshot AI KimiMoonshot AI Kimi快手指数快手指数Latent SpaceLatent SpaceLessWrongLessWrongLil'Log (Lilian Weng)Lil'Log (Lilian Weng)Linux.doLinux.doLMSYS BlogLMSYS BlogLobstersLobsters英雄联盟英雄联盟LWN.netLWN.netMacRumorsMacRumors杂志与人文网站杂志与人文网站Make: MagazineMake: Magazine猫眼电影猫眼电影Maven CentralMaven CentralMediumMediumMeituan LongCatMeituan LongCatMeta AI BlogMeta AI BlogMeta EngineeringMeta EngineeringMidjourney UpdatesMidjourney UpdatesMiniMaxMiniMaxMIT News · RoboticsMIT News · RoboticsMIT Technology ReviewMIT Technology Review米游社米游社Mozilla.ai BlogMozilla.ai BlogNatureNatureNature · Machine learningNature · Machine learning每经网与经济观察网每经网与经济观察网网易云音乐网易云音乐网易新闻网易新闻Neuroscience NewsNeuroscience NewsNew Atlas · RoboticsNew Atlas · RoboticsThe New YorkerThe New YorkerNews Hacker|极客洞察News Hacker|极客洞察水木社区水木社区NGANGA9to5Mac9to5MacNodeSeekNodeSeek牛客牛客NuGetNuGetNVIDIANVIDIA纽约时报纽约时报OEISOEISOne Useful ThingOne Useful ThingOpen Robotics BlogOpen Robotics BlogOpenAIOpenAIOpenAlexOpenAlexOpenReviewOpenReviewOpenRouterOpenRouterPackagistPackagist远景论坛远景论坛人民网人民网PhoronixPhoronixPhys.orgPhys.orgPixivPixivPlanet ROSPlanet ROS吾爱破解吾爱破解PoliticoPolitico时政媒体时政媒体中国电建阳光采购网中国电建阳光采购网专业社区专业社区Product HuntProduct HuntPubMedPubMedPyPIPyPIQoderQoderQQ音乐QQ音乐腾讯视频腾讯视频腾讯视频热搜榜腾讯视频热搜榜Quanta MagazineQuanta MagazineQwenQwenReutersReutersRFC EditorRFC EditorRobohubRobohubRobotics & Automation NewsRobotics & Automation NewsRobotics TomorrowRobotics TomorrowROS DiscourseROS Discourse阮一峰的网络日志阮一峰的网络日志RubyGemsRubyGemsRunwayRunwaySam AltmanSam AltmanScienceAlertScienceAlertScience 杂志Science 杂志Ahead of AIAhead of AI安全与破解社区安全与破解社区SecurityOnlineSecurityOnlineServeTheHomeServeTheHomeServiceNow AIServiceNow AI上海媒体(上观新闻 / 解放日报)上海媒体(上观新闻 / 解放日报)Simon WillisonSimon Willison新浪新浪Singularity HubSingularity HubSky NewsSky NewsSlashdotSlashdotSmashing MagazineSmashing Magazine什么值得买什么值得买SolidotSolidotSpotifySpotify俄罗斯卫星通讯社俄罗斯卫星通讯社少数派少数派少数派少数派Stack OverflowStack OverflowStack Overflow BlogStack Overflow Blog崩坏:星穹铁道崩坏:星穹铁道证券时报网证券时报网SteamSteamSubstackSubstackSunoSunoSuno BlogSuno BlogSynced ReviewSynced Review深圳证券交易所深圳证券交易所淘宝逛一逛淘宝逛一逛技术博客与 newsletter feed技术博客与 newsletter feedTechCrunchTechCrunchTech Xplore · RoboticsTech Xplore · Robotics腾讯热点腾讯热点腾讯新闻栏目腾讯新闻栏目The AtlanticThe AtlanticThe DecoderThe DecoderThe GradientThe GradientThe GuardianThe GuardianThe RegisterThe RegisterThe VergeThe Verge澎湃新闻澎湃新闻百度贴吧百度贴吧TikTokTikTokTindie BlogTindie BlogTomer TunguzTomer TunguzTom's HardwareTom's Hardware今日头条今日头条Transformer CircuitsTransformer Circuits旅行游记旅行游记TVmazeTVmaze优设网优设网UiverseUiverseV2EXV2EXVentureBeat · AIVentureBeat · AI万方数据万方数据中央气象台中央气象台WikidataWikidataWikipediaWikipediaWiredWired人人都是产品经理人人都是产品经理The Wall Street JournalThe Wall Street JournalxAI NewsxAI News小鹅通小鹅通喜马拉雅喜马拉雅喜马拉雅喜马拉雅新华网新华网YahooYahooYahoo FinanceYahoo Finance第一财经与 21 财经第一财经与 21 财经YollomiYollomi有道精品课有道精品课YouTubeYouTube游研社游研社知乎日报知乎日报知乎知乎Zhipu AI ResearchZhipu AI Research财新数据通财新数据通央视新闻联播央视新闻联播CelesTrakCelesTrakCISA KEVCISA KEV财联社财联社CoinDeskCoinDesk法布财经法布财经富途牛牛富途牛牛格隆汇格隆汇HDXHDX和讯网和讯网Investing.comInvesting.com界面新闻界面新闻金十数据金十数据21财经21财经金融界金融界Launch Library 2Launch Library 2MKTNews · 快讯MKTNews · 快讯NASA EONETNASA EONETNOAA/NWSNOAA/NWSNVDNVDopenFDAopenFDAOSV.devOSV.dev量子位 · 具身智能量子位 · 具身智能新浪财经新浪财经Spaceflight NewsSpaceflight NewsTelegram OSINTTelegram OSINT腾讯混元研究腾讯混元研究USAspendingUSAspendingUSGS EarthquakesUSGS EarthquakesWHO 疫情通报WHO 疫情通报选股宝选股宝第一财经第一财经

实时热搜

  • 0187
    Long-WAM: Scaling the Context of World-Action Models
    Real-time robot control demands enough visual history to infer motion and task progress, but processing that history can delay action. We present Long-WAM, a model-system framework for scaling the context of causal world-action models under real-time control constraints. Our central finding is that access to history is not the same as using it: longer histories pay off far more when the video foundation is pretrained autoregressively (AR). We first learn causal prediction from robot and egocentric videos without action labels, then preserve this history-to-future structure during world-action adaptation. On RoboCasa GR-1, increasing context from 0.0 to 19.2 seconds raises success from 63.3% to 78.7%, whereas a bidirectionally pretrained initialization shows no net gain; robot-domain AR pretraining further raises peak success on GR-1 and LIBERO-Long. Long-WAM also achieves the best results among compared methods on LIBERO-Long, RoboTwin 2.0, and DOMINO. Streaming observation encoding, asynchronous execution, and hardware-specific acceleration enable deployment on RTX 5090, DGX Spark, and Jetson AGX Thor without dropping future prediction; on RTX 5090, each action chunk, including future-video latent prediction, takes 107.4 ms. Real-time deployment on Unitree G1 and YAM supports dynamic and long-horizon manipulation, including 95% success on dynamic cup stacking, where Pi0.5 and Fast-WAM succeed in none of 20 trials. As a memory-informed executor, Long-WAM also complements higher-level planning in composite tasks.Wei Huang, Bohan Zhang, Chenzhi Liu et al.
  • 0262
    GRACE: Generation-aware latent compression for efficient video generation
    Highly compressed video autoencoders offer an effective way to accelerate video diffusion models, as the Diffusion Transformer (DiT) operates on far fewer tokens. However, such autoencoders are challenging to train, since a higher compression ratio degrades reconstruction quality and recovering it requires more channels, which is known to slow the convergence of the DiT. The compressed latent also differs from the one the DiT was trained on, so the pretrained DiT must be either retrained from scratch or adapted at considerable cost. Compressing the autoencoder the DiT was trained with appears to preserve compatibility, yet optimizing it for reconstruction alone still shifts the latent away from the distribution the DiT has learned. To address this, we propose Generation-Aware Latent Compression for Efficient Video Generation (GRACE), a two-stage framework that compresses a pretrained video autoencoder while keeping it compatible with the pretrained DiT. Specifically, we keep a frozen base latent from the pretrained encoder and learn a residual latent for the information lost under stronger compression, while aligning the compressed latent with the pretrained latent in the feature space of the frozen DiT so that the autoencoder is optimized for generation. We then adapt the DiT with lightweight fine-tuning and asymmetric denoising, where the base is denoised ahead of the residual. GRACE reduces the token count of Wan2.1-I2V-14B by 8x and its latency by 11.1x at 480x832x81, while matching the generation quality of the pretrained pipeline before compression on VBench.Jiyoung Kim, Paul Hyunbin Cho, Jisu Nam et al.
  • 0350
    SGF+: Decoupling Gradient Flows for Autoregressive Video Generation
    Autoregressive video generation requires denoising the current frames while writing their key-value representations as context for future predictions. However, these two roles typically share parameters, and we find that their gradients exhibit distinct patterns and systematic negative alignment, hindering the joint optimization of visual quality and temporal consistency. We introduce Self Gradient Forcing Plus (SGF+), which assigns separate parameters to context writing and denoising while preserving their interaction through causal attention. Both roles are jointly optimized using the original generation objective without auxiliary losses, with context writing supervised through its contribution to future predictions. This simple change improves visual quality and long-horizon consistency over the evaluated baselines in both framewise and chunkwise generation, without additional video training data or a longer training horizon. Trained on only 5s rollouts, SGF+ supports continuous generation for up to 24 hours without long-video fine-tuning. These results highlight role-specific parameterization as an effective design principle for high-quality autoregressive video generation and native long-horizon extrapolation.Zihan Su, Junhao Zhuang, Yaowei Li et al.
  • 0450
    UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation
    Dense visual text requires image generators to reproduce long strings across multiple regions with correct placement and legibility. As short-string rendering improves, evaluation must test sustained performance across more demanding scenes. We introduce UltraText Bench, a bilingual benchmark for prompt-only generation of dense visual text. It contains 432 prompts spanning 24 real-world scene categories and three difficulty levels, split equally between English and Chinese. Each human-reviewed prompt supplies exact strings for four to twelve text regions, paired with structured references for their content, placement, and visual attributes. We use the Q-Judger vision-language model to assess each image against the complete reference, reporting text fidelity, text clarity, spatial quality, and scene quality. Across 24 model configurations, these dimensions reveal different strengths: Z-Image-Turbo gains 3.81 clarity points over Z-Image-Base while losing 14.76 fidelity points under the reported settings. Performance also varies with workload; Qwen-Image-2512's English composite falls from 86.50 at L1 to 42.86 at L3. Ten participants took part in human evaluation of the automatic scores. Repository: https://github.com/LINs-lab/UltraText_Bench.Deyuan Liu, Yihao Hu, Jingxuan Zhang et al.
  • 0535
    Tetris3D: 3D Scene Generation With Objects That Fit Together
    We propose Tetris3D, a generative framework for single-image 3D scene reconstruction that recovers objects which are physically and geometrically coherent as a scene. Existing methods often generate objects independently or couple them implicitly, providing limited guidance for ensuring fine-grained spatial compatibility between neighboring objects that interact with one another. To address this, we explicitly condition the generation of each object on the geometry of surrounding objects and their physical relationships, guiding its shape and pose to remain geometrically and physically plausible within the scene. Moreover, we introduce ComOb, a physics simulation-based dataset of 1.2M scenes featuring physical interactions across diverse object categories, with per-object meshes and pairwise physical relation annotations. Comprehensive experiments on synthetic and realworld scenes show that Tetris3D recovers coherent object shapes and poses even when interacting regions are occluded, and achieves state-of-the-art performance in both generation quality and physical stability.Jaeyeong Kim, Jinhyuk Jang, Jongmin Lee et al.
  • 0634
    RunningTab: Direct Workspace Interaction with Environment-Side Tabs
    Much knowledge work produces new deliverables from files a workspace already holds, and LLM agents are beginning to take such work over. Through direct corpus interaction, an agent can search and read any of those files from a terminal with no indexing, and producing a deliverable from many of them in this way is what we call direct workspace interaction (DWI). Reaching the files, however, is only half the task: nothing keeps track of what the task asks for, what has been read, and what was listed but never opened, all of which slip through the context window without leaving a trace, so an agent may extract a figure and still deliver a report without it. To address this, we present RunningTab, a framework that equips direct workspace interaction with an environment-side tab: a per-task record of what the task still owes, kept by the environment alongside the agent. Specifically, the agent adds its requirements, while the environment records every file read as an excerpt with its provenance and every listed but unopened file as a candidate; the agent can then see each requirement beside its best-matching excerpts and top unopened candidates, resolve it against matching content or set it aside with a reason, and, should it try to finish with requirements still open, receive them in a finish check. We validate RunningTab on three benchmarks with three LLMs, where it consistently outperforms plain DWI and baselines that keep the record in the model, while its tab usually holds the values a deliverable needs once seen.Jinheon Baek, Soyeong Jeong, Yumin Choi et al.
  • 0729
    RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments
    General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world. To investigate this, we introduce RobotWorld, a challenging simulation testbed for robot use: turning instructions and observations into physical task execution through robot interfaces. Its 84 tasks span manipulation, mobile manipulation, locomotion, driving, and aerial control, with explicit interaction budgets and executable success checks. By analysing task outcomes alongside execution traces, we identify both the capabilities that transfer and the gaps that prevent reliable completion. Furthermore, we find that current agents can construct sophisticated perception and control workflows, including image segmentation, camera calibration, spatial estimation, and dynamics-based computation. These capabilities, however, do not consistently compose into successful behaviour: agents lose task-relevant object states despite reaching commanded poses, fail to correct ineffective actions, recover too late, or mistake unfinished tasks for completion. This uneven transfer also differs across models: Astra succeeds more often on spatial and constrained-contact goals, whereas Opus 5.5 succeeds more often on continuous-balance and timed-interaction goals. By linking these outcomes to execution behaviour, RobotWorld provides both a rigorous proving ground and an empirical account of the remaining capability gaps, thereby establishing concrete targets for training and designing more reliable physical-world agents.Zhiqin Yang, Chenxin Li, Xiaomeng Hu et al.
  • 0829
    Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position
    The architectural design of Large Language Models (LLMs) is shifting from traditional full-attention-only models to hybrid models, which combine different attention modules to improve long-context efficiency and performance in length extrapolation and context extension. To explain why hybrid models work and how to design them better, we propose Mechanics of Long-Context Hybrid Models. As Part 1.1 of this series, we begin with hybrids of full attention and either sliding-window attention (SWA) or gated variants of linear attention (LA), represented by GLA and GDN. We first observe a Seesaw Effect in Context Extension: LA hybrids benefit more from long-context continual pretraining, whereas SWA hybrids perform better under length extrapolation. We attribute this behavior to differences in the positional inductive biases induced by these attention mechanisms. We find that SWA hybrids suffer from a Short-Context Learning Trap, Short-Window Weariness, and Long-Window Laziness, and require extended windows to enhance performance in continual long-context pretraining. For LA hybrids, we summarize the Matthew Effect of Hybrid Position Extrapolation and propose Sliding-Window Linear Attention, achieving 16times training-free length extrapolation while maintaining 100\% accuracy on NIAH-SK1 in 64k context length.Xiaoran Liu, Ziwei He, Xipeng Qiu
  • 0919
    RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning
    Recently, approaches that leverage human video datasets for robot policy training have become increasingly prevalent. However, most existing hand trackers regress pose from cropped frames with limited priors on hand motion and object interaction, resulting in inaccurate and physically inconsistent estimates. Moreover, the lack of physical cues, e.g., contact and force, limits the use of human videos for robot policy training. To this end, we propose RLHND, a video foundation model-based hand tracking model that jointly estimates hand pose and realistic tactile information from monocular egocentric videos. RLHND turns the pre-trained Cosmos 3 video diffusion backbone into a deterministic clip-level feature extractor via clean-latent conditioning, carrying its learned priors on hand motion and hand-object interaction into tracking. For pose estimation, RLHND (i) predicts hand poses with anatomically plausible joint angles and (ii) enables optional conditioning on the shape parameter to maintain consistent hand shape within the same video and even across videos recorded by the same actor. For tactile estimation, a separate tactile expert stream, trained with the pose stream frozen, predicts dense contact and force over the hand surface. We further adopt LBS-based feature spreading to enable vertex-wise feature extraction without costly per-vertex attention. RLHND achieves state-of-the-art performance across various benchmark datasets for pose estimation, while also achieving state-of-the-art performance in contact and force estimation. Moreover, we demonstrate the utility of RLHND for robot learning through retargeting results and real-world robot experiments. The code will be publicly available at https://seungjun-moon.github.io/rlhnd/.Seungjun Moon, Subin Jeon, Sangwoo Kim et al.
  • 1018
    Inverting Multi-Vector Visual Document Indices
    Prevailing multi-vector visual document retrievers store each page as about a thousand patch vectors, often in vector databases run by a third party. Since no one can read a page from its vectors, this index is easily treated as less sensitive than the page. However, because the index keeps one vector per patch in raster order, and each vector is computed by a vision-language model pre-trained to read documents, we hypothesize that whoever runs or breaches the store can reproduce a page from its index alone. We frame inversion as conditional document image generation and infer from the vectors what the attack needs: the encoder, the page shape and, for shuffled vectors, their order. On the ViDoRe v3 benchmark, pages inverted from raw indices recover 47% of the words and 45% of the sensitive tokens. Used as queries against the stored indices, they rank their source page first 98.4% of the time. We test two cheap protections, token pooling and shuffling, which both cut word recall to about 8%. A model that restores the order of a shuffled index raises the share of source pages ranked first from 3.8% to 93.5%, while inverting a pooled index remains open. To test generalisation, we apply the same attack unchanged to another multi-vector retriever: its inverted pages still rank their source page first 70.2% of the time, though its word recall stays below a nearest-neighbour baseline. Multi-vector visual document retrievers are therefore vulnerable to inversion through their stored index, which should be protected like the documents it encodes.Zhuchenyang Liu, Yao Zhang, Yu Xiao
  • 1116
    AdSpark: A Large-Scale Dataset and Benchmark for Product-Centric Advertisement Video Generation
    Product-centric advertisement video generation aims to create promotional videos that preserve fine-grained product identity while presenting selling points through coherent multi-shot narratives. However, this emerging task remains underexplored due to the lack of large-scale advertisement-specific datasets and comprehensive evaluation frameworks. To address this gap, we introduce AdSpark, a large-scale dataset and benchmark for product-centric advertisement video generation, based on data from a major e-commerce platform. AdSpark-300K contains approximately 300K reference image--prompt--video triplets, comprising a real-world subset and a synthetic subset. Each sample provides structured advertisement annotations, including product identity annotations, selling-point descriptions, creative plans, and aligned audio scripts, enabling models to learn product preservation and advertisement-oriented visual storytelling. We further propose AdSpark-Bench, a diagnostic benchmark that evaluates generated advertisements across six dimensions, including visual quality, product fidelity, instruction adherence, temporal coherence, audio alignment, and advertisement effectiveness. Based on AdSpark-Bench, we evaluate representative models, revealing key challenges in product preservation, multi-shot storytelling, and selling-point visualization. Experiments with AdSpark-300K-finetuned models further validate the effectiveness of our dataset. AdSpark provides a unified dataset and benchmark for future research, and we will release the dataset upon acceptance.Zhifei Yang, Zhao Jiang, Keyang Lu et al.
  • 1216
    From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery
    Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--cost or accuracy--latency Pareto frontier. Yet user requirements are multidimensional: users may specify accuracy, latency, and inference-cost requirements jointly, and different requirements can favor different controllers. We formulate Personalized Test-Time Scaling as discovering executable controllers that maximize the joint satisfaction rate of user-specific requirements. To reduce the overhead of repeated policy discovery for new user profiles, we propose PersonTTS, an amortized agentic policy-discovery framework that reuses prior search experience through requirement-matched controller initialization and source-distilled procedural guidance, while retaining target-profile evaluation for every candidate. Experiments on AIME and HMMT show that PersonTTS substantially outperforms strong TTS baselines in joint requirement satisfaction on unseen user profiles and held-out problems. Under the same candidate-evaluation budget, cross-user experience reuse further improves policy quality while substantially reducing discovery-agent time and cost.Xinglin Wang, Zishen Liu, Tong Zheng et al.
  • 1316
    WebFovea: When the Model Is Right but the Click Is Wrong -- Reliable Round Trips for Vision-Based Web Agents on Live Websites
    We present WebFovea, a vision-based web agent that placed 2nd in the WebRetriever Challenge 2026 with a final score of 57.0 out of 100. The challenge evaluates agents end to end on Protocol III of the WebRetriever benchmark (arXiv:2607.06118): starting from an entry URL on a live website, the agent must operate the site's own interface and return a verifiable answer. A capable multimodal large language model (LLM) is necessary for this, but not sufficient. The model's decisions reach the browser through the harness, the code between the model and the page. At every step, four things must go right: the model's reply must be parsed into the intended action, the action must take effect on the page, the result must be reported back accurately, and the model must be shown the information it needs. On real websites, many of the failures we observed occurred at one of these four stages rather than in the model's reasoning. A coordinate-space mismatch placed every click at 3/4 of its intended coordinates; actions on native dropdowns, inside iframes, and in text boxes failed silently; and self-generated chat-template tokens contaminated 4.9% of task episodes. WebFovea hardens each stage and surrounds the loop with guardrails that keep the agent within the rules and its budget. The four-stage view does not depend on the model, although some individual fixes do. Because we used the same model in all four submissions, the rise of our official hidden-set score from 31.0 to 57.0 reflects changes to the harness, up to run-to-run variance on live sites. We describe the design, the evidence for each component (including negative results), a failure analysis, the limitations, and a roadmap that includes routing different steps to different models. Code is available at https://github.com/jianganghan/WebFovea.Jiangang Han
  • 1414
    Q-Learning with Scalar Adjoint Matching
    Flow policies capture rich and diverse action distributions, and fine-tuning them with off-policy RL to improve beyond the demonstrations has drawn growing interest. However, fine-tuning a flow policy against a learned value function is not trivial, because the policy generates its action over many flow steps. Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size. We observe that the batch-averaged velocity Jacobian of pretrained flow policies concentrates on its diagonal. Motivated by this finding, we derive a closed-form scalar adjoint that scales the value gradient at the final action by the flow time, eliminating the per-step vector--Jacobian products. We further find that controlling the critic's value at policy-generated actions is particularly important under the scalar adjoint. Based on these findings, we propose Q-learning with Scalar Adjoint Matching (SQAM), which combines the scalar adjoint with a value penalty at those actions. SQAM's gains concentrate on the four hardest OGBench domains, where its success rate exceeds that of the strongest baseline in each domain by 18 to 35 percentage points. To test whether SQAM extends to large pretrained policies, we also fine-tune a vision-language-action policy on a real bimanual robot. SQAM improves over supervised fine-tuning on all three tasks.Yonghoon Dong, Minsung Yoon, Jaehyuk Kim et al.
  • 1512
    QuadTok: Quadtree Visual Tokenizer for Autoregressive Image Generation
    We introduce QuadTok, a novel framework for visual tokenization and autoregressive image generation. Compared to traditional approaches using 2D grids or 1D token sequences, we propose a hierarchical quadtree structure, bridging the gap between 2D spatial binding and 1D sequence-level flexibility. The QuadTok tokenizer dynamically allocates representational capacity to visually intricate areas while leaving homogeneous regions at a coarse resolution. Compared with a fixed 256-token grid, our ImageNet-trained tokenizer saves approximately 10% of tokens on ImageNet and 9% when transferred zero-shot to the COCO dataset, while maintaining comparable reconstruction fidelity. Furthermore, the natural causality introduced by the tree structure seamlessly enables autoregressive image generation. Conditioned on a quadtree topology supplied before generation, our 947M GPT-style generative model achieves a 2.08 gFID on the ImageNet 256 times 256 benchmark. Additionally, leveraging the strong spatial correlation preserved by the quadtree structure, the QuadTok generator enables zero-shot spatially controlled image generation capabilities. Code: https://github.com/myc634/QuadTok.Yucheng Mao, Zeyuan Chen, Xiaojun Shan et al.
  • 169
    MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
    Training and evaluating interactive language agents typically requires rich user interactions, yet collecting human feedback is expensive and difficult to scale. Simulated users offer a scalable alternative, but they must both resemble real user behavior and provide useful learning experiences for agents. In contrast, most agent-training frameworks rely on off-the-shelf assistant LLMs, whose helpfulness can make them overly cooperative, explicit, and behaviorally homogeneous compared with real users. We introduce MIMESIS, a purpose-built user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns derived from real user interactions. Empirically, our 9B model achieves a SOUL-Index of 65.7, surpassing the strongest frontier model. Compared with Claude-Opus-5, the strongest baseline on RealUserSim and SimulatorArena, MIMESIS improves behavioral fidelity by 13.4 points and reduces Turing distance by 3.6 points, respectively. We then freeze the simulator and train an agent by interacting with the frozen simulator using multi-turn reinforcement learning. Across eight environments, training with MIMESIS yields better agent performance than training with GPT-5.5 under all nine unseen user simulators, demonstrating stronger generalization to new user simulators. Moreover, we propose Coached On-Policy Self-Distillation (CSD), which leverages simulator-generated private reasoning traces and subsequent utterances as feedback on how well the agent addresses user needs. A coach converts this information into concise coaching notes that describe how the agent can better anticipate user needs and adapt its behavior over the course of an interaction. CSD turns this feedback into dense, token-level supervision beyond sparse task rewards, yielding further gains across all nine evaluation user models.Hoang Phan, Dat Huynh, Andrey Zhmoginov et al.
  • 177
    SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles
    Memory-augmented reinforcement learning strengthens LLM agents' ability to solve complex long-horizon tasks. Skills are one such form of memory, pairing instructions with an applicability condition over task types. However, retaining every skill indiscriminately as the policy improves lets obsolete or harmful entries accumulate and mislead the agent. We propose SkillForge, an agentic RL method that compiles and evolves the skill library through a fitness-driven skill lifecycle of trial, active, stable, and retired states, so that the skills and the model co-evolve throughout training. A pre-RL evaluation phase first uses the base model's own rollouts to pre-retire low-fitness skills, yielding a filtered library that then seeds supervised fine-tuning. Reinforcement learning takes over from this checkpoint, and at each iteration selective retirement, stabilization, and LLM-guided mutation continue to forge the skill library alongside policy optimization. Across multiple interactive agent benchmarks, SkillForge achieves the highest aggregate success rate, delivering up to 7.8% relative improvement over the strongest baseline while keeping the skill library compact throughout training. We introduce SkillFurnace, a dataset of 5k+ annotated records bundling retirement-filtered SFT trajectories, evolved skill libraries with fitness annotations, and retirement events with human-annotated failure categories to support research on skill quality and lifecycle management.Yuyao Ge, Yiwei Wang, Yuchen He et al.
  • 186
    Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts
    Multi-teacher on-policy distillation (MOPD) is used in two settings. In common-domain composition, several teachers score each student rollout from one prompt domain and their signals form a single target; in routed-domain distillation, prompts from different domains are assigned to the corresponding specialist. Both settings usually transfer each teacher's endpoint policy, which mixes what post-training changed with preferences inherited from the teacher's base. We introduce Δ-MOPD, which transfers each teacher's teacher-minus-base logit shift re-anchored at the student's frozen initialization, and compare it with endpoint supervision in both settings while holding teacher selection fixed. We first expose the mechanism that impedes endpoint transfer: inherited base pull can exceed the post-training shift. Removing it reduces the teacher-term norm ratio and target--student KL. Across our experiments, the results suggest that shift targets are particularly useful when teacher signals are combined at a state. With three composed teachers, Δ-MOPD exceeds endpoint composition by 4.11 Math and 1.95 five-benchmark points; with two, it matches endpoint accuracy. Under phased routing, it achieves higher mean performance in both phase orders and reduces the observed order gap from 10.50 to 6.42 points. Under interleaved routing, where each update involves one teacher, the two targets perform comparably. The phased results provide supporting evidence that the benefit may extend to signals accumulated across training phases. Target construction is thus an independent design axis in MOPD, complementary to teacher selection.Hejian Sang, Zhengze Zhou, Shayan Mohajer Hamidi et al.
  • 196
    Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
    Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters. We test this claim along both routes to a pixel-space backbone. We pretrain Iris-3B, a 3B-parameter pixel-space text-to-image transformer, from scratch through a 256to512to1024 curriculum, after first ablating the prediction target and representation alignment at 256^2 to decide what to scale. We also convert a pretrained latent model, FLUX.2 Klein base 4B, to pixel space. We fine-tune both families for monocular depth estimation and for image restoration/super-resolution. We find no significant improvement from using a pixel-space generative prior. Fine-tuned for depth with one matched direct-regression recipe, Iris-3B is level with the latent FLUX.2 Klein and the converted pixel FLUX.2 Klein falls behind it, and on 4times DIV2K restoration neither pixel model beats a latent FLUX.2 Klein fine-tune, the converted one trailing it slightly. We document the recipes, the failure modes and the remaining confounds behind this negative result. Nevertheless, Iris-3B shows that pixel-space pretraining with the pixel-transformer (PiT) head of PixelDiT scales to 3B parameters and to text-to-image quality competitive with latent models, matching Qwen-Image on OneIG under the official evaluators at 1024^2. We release its weights and training code in the hope that they help pave the way for further work on pixel-space generation.Hanqiu Li Cai, Chema Garabito
  • 206
    We Query, Therefore We Compute: On Oracle Computation beyond the Machine, with an Application to Agents
    Agentic systems use large language models (LLMs) to carry out concrete tasks. Prior work often borrows abstractions such as scheduling, caching or isolation piecemeal from operating systems, so the mechanisms it builds share little common ground, and the shared view of the two forms of agentic system, Workflows and Agents, is limited. We construct an abstract machine that provides both. We treat the LLM as an Oracle and extend a two-stack pushdown automaton with one instruction, which hands the Oracle a whole stack as its query and appends the answer to that same stack. The machine thus performs two computations, the Oracle's and a Turing-complete one that we call the Priestess. A stack that the program only appends to grows autoregressively, as an agent's context does. Two symmetry breakings, S in storage and T in transitions, make a Priestess program the operating system of the programs the Oracle runs, and produce the Agent and the Workflow as the two placements of a task's program. For internally autoregressive Oracles, the two computations synchronize at the end of every answer under certain conditions, and through that synchronization we model caching and analyse scheduling. No guarantee that holds for every Oracle can fix which content crosses between the two computations, but such a guarantee does fix the boundary itself. The construction V fits the machine to a von Neumann computer. To show that it is realizable, we propose ArchNights, an extended RISC-V ISA and a Linux-style operating system implementing the machine by design. ArchNights-SE runs on gem5 as a computer system, becomes an agentic system when it runs an LLM as the Oracle, and will be open source. Agentic systems can then be designed as computer systems are. With a foundation built and a unified view, future work can share invariants and bounds, each with its conditions.Kefan Liu, Fengning Ou, Yelin Luo et al.
  • 017246
    余红旧事
    一片广场上发生一连串荒诞命案——冻成冰雕的酒鬼,惨被断肢的壮汉,人间蒸发的地头蛇……看似毫无关联,实则命运环环相扣,情感彼此纠葛。
  • 024487
    黑岛监狱
    凌东国刑警廉明为了破获边境人口失踪案,也前往黑岛监狱卧底查案。南岭国退役特种兵文傍为了营救和自己相依为命的雅醇母女,被迫进入黑岛监狱服刑。狱中二人从互相敌对到彼此扶持,互助协作,勇敢对抗黑帮势力的同时,还要机智化解特殊身份给他们带来的重重危机。就在二人逐步揭开监狱犯罪真相的同时,一场更大的危险正在渐渐靠近。
  • 033821
    深渊无间
    一篇名为《深渊》的推理网文悄然上线,打破了保守小城多年来的平静,文中诸多情节与警方未曾公布的多年前悬案案情有着惊人的相似。热血正义的新警李成,与多方嫌疑人,一次次上演高智对弈。最终李成拨开迷雾,侦查出掩藏在令人扼腕的亲情和友情之下的真相。
  • 042894
    凛冬下的罪恶
    凛冬,一精神病女子拦车报案,称丈夫杀人,刑警沈栋梁吴红兵由此揭开系列碎尸案真相。然而风浪未平,储蓄所抢劫杀人案,少女失踪案,流窜抢车案接连发生,沈栋梁与吴红兵追凶之际,竟牵出改变二人命运的人性悲剧。
  • 052515
    昨夜将至
    面馆老板韩栋与妻子林美月看似安稳的日常之下,各自埋藏着不愿被人知晓的过往。林美月曾经的身份被旧识要挟勒索,平静生活被骤然打破;韩栋尘封二十年的秘密,也随着一场蓄意的复仇逐渐浮出水面。旧友步步紧逼,夫妻二人被卷入层层交织的危机当中。多年前的遗憾与过错、旧日姐妹间的纠葛接连爆发,多方势力相互拉扯。为守护自己的小家,夫妻俩从被动周旋开始奋力反击,在迷雾重重的恩怨里,直面所有过往造成的困局。
  • 062318
    她的罪名
    广南市出租屋发生恶性双尸案,夜总会女子李娜琳达惨遭虐杀。刑警林威与经验老到的马国栋联手查案,从现场细节推翻劫财劫色判断,锁定失踪服务员曹爱媛及其背后团伙。经查,曹爱媛因长期受死者轻视羞辱,心怀怨恨,联合华卫红等人设局行凶。警方循线追踪,揭开团伙谋财害命、内讧逃亡的真相。案件跨越三年,曹爱媛隐姓埋名再犯命案,最终被林威抓获归案。老刑警马国栋带病追凶,壮烈牺牲。故事以连环命案为引,剖开原生家庭、人性善恶与底层挣扎,展现刑警坚守正义、追凶到底的使命担当。
  • 072284
    蝉
    生父是绑架犯,继父是刑警的刑辩律师(钟楚曦 饰),多年前遭人报复痛失妻儿的大法官(吴镇宇 饰),善用人心,游走灰色地带的邪魅律政精英(郑云龙 饰),三个孤独的陌生人在硬核案件的刑事辩护中互相博弈,却越发惺惺相惜,是对手也似亲人,螳螂捕蝉,黄雀在后,三人间,互为猎物,亦互为猎手,究竟谁才是黄雀,谁又是那只蛰伏十二年,只为自由鸣叫,收获一季响亮夏天的蝉?
  • 082202
    漂白
    《漂白》融合犯罪、刑侦、心理学等多重元素谋篇布局,讲述了刑警队长彭兆林对正义初心执着坚守,在明与暗的角力中不断追逐,历经十年跨越千里,终将极端偏执与残忍弑杀的罪犯绳之以法的故事。实力演员郭京飞、王千源、赵今麦主演,曹凯导演执导,陈枰编剧,展现了人民警察流淌在血液里的信念和坚持,邪不压正,正义终将得到伸张!
  • 092068
    除恶
    该项目改编自雷米的小说《老男孩》,讲述了一个偏僻闭塞的沿海小镇,被一袋消失的毒品打破沉寂,一场“自杀式拯救”正式拉开序幕。一个是意外卷入毒品交易现场的女刑警胡文静,一个是为救女儿不计代价的普通父亲程恳,一个是一心想要逃离小镇生活的野心家李晓雅,一个是被毒品吞噬而走向异途的姐弟王萍王安。欲望与贪念交织成的茧细细密密困住了每一个人,无人能够全身而退。剧中沉浸式塑造小镇生态,力求打造小而美、充满生活感的高水准缉毒剧。
  • 102040
    暗金
    警校学生燕鸣宇没能入警,又因参与地下拳赛被开除。经侦队长陈昌杰安排他卧底生父麦忠伟的天曜集团,调查812洗钱案,追查母亲死亡真相。燕鸣宇随后得知,坎坷遭遇全是组织布置的卧底考验。他逐步获取信任,查清集团勾结黑恶势力借拍卖会洗钱,并成功执掌公司。收网前夜内鬼出手,燕鸣宇卧底身份暴露,麦忠伟舍身救子却遭章俊生杀害,燕鸣宇被诬陷通缉,队友重伤。燕鸣宇假意投靠高官章致远,依靠加密货币掌握15亿赃款与罪证。最终内外联动一网打尽所有罪犯,陈年命案水落石出燕鸣宇沉冤得雪,重返警队。
  • 112030
    深渊
    在铜江市,短短三个月内连续爆发了三起恶性案件:先是游客在城郊大黑山发现了江中漂浮的残臂;随后,大学生路阳在家中不明原因死亡;紧接着,市青年企业家叶启明遭遇车祸不幸身亡。市局刑侦科的女刑警蒋梅与刚从省厅调来的沈峰奉命侦破这一系列案件。随着调查的艰难推进,案件背后残忍的杀戮、深藏的爱恨情仇以及错综复杂的关系渐渐显现,而血案背后的缘由也折射出物欲诱惑与扭曲心灵之间无法割裂的因果联系。
  • 121953
    有罪之身
    一次失手误杀,一场牵扯十三条人命的爆炸案,两个看似毫不相干的案子,因为三个年轻人交织在一起,跨越十年的调查与救赎,终于引出背后的惊人秘密。
  • 131869
    三大队
    讲述一次审讯意外,三大队刑警程兵入狱服刑,队友受牵连脱警、降职,曾经的警界精英三大队分崩离析。十年牢狱,程兵重获自由,失去一切,而案件的犯罪嫌疑人王大勇依旧在逃…… 穿一天警服,终身是正义。三大队需要交代,不甘化作执着,利刃再次出鞘,程兵和三大队的兄弟重新集结踏上追凶之路,在孤独和漫长的旅途中配合警方千里追凶,也在这苦行僧一样的历程中重新找到人生的坐标和生命的意义。 本片根据原载于“网易人间” 作者深蓝 的《请转告局长,三大队任务完成》改编。
  • 141855
    不可告人
    一次既成功又失败的抓捕行动,改变了所有人的命运,痛苦中的等待意味着从未放弃,而那些不可告人的罪恶也在伺机重来,这一次,是正与邪之间最终的较量!
  • 151762
    赦罪
    沈庆明怀孕的妻子被栾大海醉驾撞死,栾大海花钱买通了别人顶罪。沈庆明冲动之下打伤了栾大海入狱。出狱后,他处心积虑想杀死栾大海报仇,却因良知与突如其来的爱情,在即将下手的时刻放弃。不料,栾大海竟然惨死,所有迹象都指向沈庆明,而当真凶浮出水面时,他才愕然发现,真凶竟然是他无论如何也想不到的人……
  • 161735
    错位
    讲述了刑警姜光明(马伊琍 饰)和石落(高至霆 饰)在调查一起案件时,偶然发现作家顾己鸣(佟大为 饰)的小说中所描绘的犯罪现场与自己正在调查的案发现场离奇重合......虚构与现实交错,小说的出现为二人的追查提供了新的方向,但也将他们引向了更深的迷局......案中案,谜中谜,隐藏在幕后的凶手究竟是谁?
  • 171625
    猎罪图鉴
    该剧讲述了因一起尘封旧案而结怨的模拟画像师沈翊和刑警队长杜城,在机缘巧合下被迫搭档,两人联手侦破多起离奇疑案,共同追踪谜底真相的故事。
  • 181607
    乌云之上
    韩青是东州市刑侦队的女警官,因为不苟言笑只问工作被送外号“不高兴”,她与搭档钟伟合作多年十分默契,然而钟伟却在跟踪一位嫌疑人后失踪,快两个月杳无音信。韩青在焦虑、痛苦中处理着其它案件,同时不放弃任何与钟伟有关的线索。新出现的命案引起韩青的高度警觉,她通过蛛丝马迹判断命案与钟伟的失踪有关,尽管调查中阻力重重,真正的嫌犯最终还是浮出水面,而钟伟到底经历了什么也一点点变得清晰起来,警方抽丝剥茧,发现了失踪、杀人、贩毒三件大案的关联处,摸到了一个犯罪集团精心组织架构的网,也发现了隐藏在这个网背后的秘密。东州市刑侦队群策群力,一举破获了全案,犯罪分子悉数落网。
  • 191588
    不眠日
    华澳市警署有一位睿智骁勇、屡破奇案的“神探”,警署内传言称没有他破不了的案件。然而,看似完美的杀人案接二连三地降临,挑战着这位“神探”的极限。案件中,凶手如同幽灵般无迹可寻,没有留下任何线索,仿佛从未来过。华澳市风云变幻,暗流涌动,每一个角落都可能隐藏着致命的危机。
  • 201565
    沉默的真相
    一起看似简单的自杀案, 背后隐藏着一个不可告人的巨大秘密; 为了揭开这个秘密, 一群人历经七载, 付出无数代价, 甚至赌上性命...一个曾有大好前途,四平八稳的检察官江阳,但因受贿贪污,坐牢三年,再次出现在公众视野里竟是他出现在行李箱里的蜷缩的尸体;运送他尸体的当地著名律师张超。此案引发全市关注,刑警严良介入负责此案,与张超正面交锋。张超承认犯罪,却在接受法院宣判时当场翻供,并拿出不在场证据,案件陷入死局。
  • 01
    AI把创新效率拉满,为什么好想法却越来越少?
    同样手握大模型,不同团队的创新输出却天差地别。AI 能够大幅提升处理效率,却容易放大人类固有的认知偏差与组织偏见。创新不全是技术问题,认清 AI 的能力边界,守住人与真实世界的连接,才是破解创新困局的关键。朱利安·德弗雷塔斯(Julian De Freitas)、阿耶莱特·伊斯雷利(Ayelet Israeli)、吉迪恩·纳韦(Gideon Nave)、阿尔乔姆·季莫申科(Artem Timoshenko)、奥利维耶·图比亚(Olivier Toubia)
  • 02
    外部空降的高管,为何很少能升任CEO?
    很多外聘高管交出出色业绩,却与 CEO 岗位失之交臂。大家常归咎于能力短板,实则容易忽视组织认同度这一隐形壁垒。继任并非只筛选能人,空降人才需要完成身份蜕变。企业的人才规划,应当把这份集体共识纳入考量。阿南德·乔希(Anand Joshi)
  • 03
    AI时代,高潜力人才与普通员工的区别,在于三种能力
    AI正在重塑雇主对新员工的技能期望。吉姆·杜塞特(Jim Doucette)、维沙尔·高尔(Vishal Gaur)
  • 04
    廉价资本时代终结,靠烧钱换增长的日子到头了
    那些严谨配置资本、有选择性地投资、并让战略与经济效益始终保持清晰联系的公司,将更有能力胜出。迈克尔·曼金斯(Michael Mankins)、马修·克鲁皮(Matthew Crupi)、周浩
  • 05
    打折不是认输,而是一种高级的赚钱艺术
    关键不在于要不要打折,而在于如何聪明地打折。拉菲·穆罕默德(Rafi Mohammed)
  • 06
    “主动道歉” 竟是错的?这项研究颠覆常识
    当服务出错时,企业第一时间道歉真的能安抚客户吗?梅森·R·詹金斯(Mason R. Jenkins)、玛丽·斯特菲尔(Mary Steffel)、保罗·W·丰贝尔(Paul W. Fombelle)
  • 07
    AI正在批量制造伪专家,这些关键信号要警惕
    传统的“思想领导力”正在失效,而一种更宝贵的能力——“思想践行力”——正在崛起。约翰·温索(John Winsor)
  • 08
    让员工真正接受任务,关键在于做好三件事
    通过提供真正的自由,在做出决策时传达确定性,并让分配过程显得合理,领导者可以培养员工的接受感,使他们能够真心投入。亚当·埃里克·格林伯格(Adam Eric Greenberg)、维基·G·莫维茨(Vicki G. Morwitz)、库尔特·P·芒茨(Kurt P. Munz)
  • 011.3万亿
    Claude Code
    1.31万亿 tokens · 1084.2万 次请求 · cli-agent
  • 021.5万亿
    Hermes Agent
    1.49万亿 tokens · 1492.1万 次请求 · personal-agent, cli-agent
  • 035920.7亿
    Kilo Code
    5920.74亿 tokens · 806.1万 次请求 · cli-agent, ide-extension
  • 044988.5亿
    pi
    4988.50亿 tokens · 443.0万 次请求 · cli-agent
  • 054751.2亿
    Cline
    4751.25亿 tokens · 494.7万 次请求 · ide-extension, cli-agent
  • 064598.7亿
    Freebuff
    4598.71亿 tokens · 327.4万 次请求
  • 074537.1亿
    Codex
    4537.14亿 tokens · 682.3万 次请求 · cli-agent
  • 082599.1亿
    codex-fleet conductor
    2599.11亿 tokens · 365.0万 次请求
  • 092440.3亿
    omp
    2440.33亿 tokens · 156.9万 次请求 · cli-agent
  • 102011.5亿
    OpenClaw
    2011.46亿 tokens · 272.8万 次请求 · personal-agent, cli-agent
  • 111580.9亿
    DeepSeek Harness
    1580.86亿 tokens · 103.9万 次请求
  • 121260.9亿
    OpenHands
    1260.90亿 tokens · 227.7万 次请求 · cli-agent
  • 13936.0亿
    Strix
    935.98亿 tokens · 75.1万 次请求 · cli-agent
  • 14783.0亿
    ISEKAI ZERO
    783.04亿 tokens · 134.6万 次请求 · game
  • 15712.3亿
    Hello Minds, powered by Ethoswarm
    712.35亿 tokens · 66.2万 次请求 · roleplay, personal-agent, creative-writing
  • 16684.0亿
    Lemonade
    683.97亿 tokens · 114.9万 次请求 · programming-app
  • 17577.7亿
    HighLevel
    577.66亿 tokens · 119.7万 次请求
  • 18504.6亿
    LangChain
    504.61亿 tokens · 266.2万 次请求
  • 19459.3亿
    Portkey AI
    459.27亿 tokens · 582.1万 次请求 · programming-app
  • 20421.4亿
    Descript
    421.44亿 tokens · 90.9万 次请求 · video-gen