
The Decoder · 实时热榜
- 01METR introduces a new metric to calculate exactly when AI agents become more expensive than humans
METR's new metric, the "expenditure horizon," puts a dollar figure on how cost-effective AI agents are at solving problems. Early results on the NanoGPT speedrun are underwhelming, the metric has blind spots, and the newest generation of models could change the picture. The article METR introduces a new metric to calculate exactly when AI agents become more expensive than humans appeared first on The Decoder .
最高第 1 名00:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时51分 - 02Delhi High Court hands OpenAI a win by rejecting major Indian news agency's copyright injunction
The Delhi High Court has handed OpenAI a major win in its copyright fight with news agency ANI. For the first time, a court has classified AI training as private use. ANI undermined its own case by citing articles published after the models were trained. The main trial is still pending. The article Delhi High Court hands OpenAI a win by rejecting major Indian news agency's copyright injunction appeared first on The Decoder .
最高第 1 名01:58 达到01:58 首次观测上榜当日结束时仍在榜累计约21小时52分 - 03Microsoft launches its own cybersecurity model MAI-Cyber-1-Flash but still depends on OpenAI for the toughest tasks
Microsoft introduces MAI-Cyber-1-Flash, a compact security model that scores 96 percent on the CyberGym benchmark when embedded in its MDASH multi-agent system. Microsoft says costs should drop by 50 percent compared to pure frontier models, since only tough cases get passed to GPT-5.4. For complex reasoning, Microsoft still relies on OpenAI. The article Microsoft launches its own cybersecurity model MAI-Cyber-1-Flash but still depends on OpenAI for the toughest tasks appeared first on The Decod
最高第 1 名03:02 达到03:02 首次观测上榜当日结束时仍在榜累计约20小时48分 - 04OpenAI says more workers are using ChatGPT to do other people's jobs
OpenAI analyzed over 800,000 work-related ChatGPT messages and found that 43.5 percent of job-specific queries involve tasks from other professions. The company calls this "task crossover." The trend is most pronounced at small businesses, where users increasingly handle specialized work without dedicated experts. The article OpenAI says more workers are using ChatGPT to do other people's jobs appeared first on The Decoder .
最高第 1 名03:18 达到03:18 首次观测上榜当日结束时仍在榜累计约20小时32分 - 05Moonshot AI releases Kimi K3 open weights and infrastructure after shaking up the frontier model race
Moonshot AI has released Kimi K3's model weights and made parts of its infrastructure open source. The Chinese model nearly matches Western frontier models such as Fable 5 and GPT-5.6 Sol on popular benchmarks, but independent tests found major gaps in cyber and math performance, possibly pointing to distillation. The article Moonshot AI releases Kimi K3 open weights and infrastructure after shaking up the frontier model race appeared first on The Decoder .
最高第 1 名03:50 达到03:50 首次观测上榜当日结束时仍在榜累计约20小时 - 06Anthropic CEO Amodei doubles down on open-weight risk stance while insisting he never called for a ban
Anthropic CEO Dario Amodei is once again warning about the risks of open AI models while insisting he has never called for a ban. He argues that authoritarian states like China could overtake the US and that open models could be misused for biological or cyberattacks. Critics say he's mostly trying to protect his own business from cheaper competition. The article Anthropic CEO Amodei doubles down on open-weight risk stance while insisting he never called for a ban appeared first on The Decoder .
最高第 1 名20:22 达到20:22 首次观测上榜当日结束时仍在榜累计约3小时28分 - 07Nvidia invests in Ilya Sutskever's AI lab, shifting SSI away from Google chips
Nvidia is pouring what it calls a "substantial" sum into Safe Superintelligence (SSI), the AI lab run by Ilya Sutskever, OpenAI's former chief scientist. The article Nvidia invests in Ilya Sutskever's AI lab, shifting SSI away from Google chips appeared first on The Decoder .
最高第 1 名21:10 达到21:10 首次观测上榜当日结束时仍在榜累计约2小时40分 - 08Taiwan detains Nvidia employee in widening China chip smuggling probe
Taiwan's prosecutors have detained an Nvidia employee in connection with the alleged illegal export of Super Micro AI servers to China, according to Bloomberg and Reuters. The article Taiwan detains Nvidia employee in widening China chip smuggling probe appeared first on The Decoder .
最高第 1 名21:26 达到21:26 首次观测上榜当日结束时仍在榜累计约2小时24分 - 09Shared Claude chats were reportedly showing up in search engines
Shared conversations with Anthropic's Claude chatbot briefly appeared in Google search results because the pages lacked a noindex tag. Users said some chats contained crypto keys and legal questions. OpenAI made the same mistake last year. The article Shared Claude chats were reportedly showing up in search engines appeared first on The Decoder .
最高第 2 名00:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时51分 - 10Cursor's agent swarm suggests cheaper models can handle most coding when frontier models plan the work
Cursor asked its upgraded agent swarm and its predecessor to rebuild SQLite in Rust using only the documentation, with no source code or internet access. Every configuration of the new system, which separates planners from workers, eventually scored 100 percent on the test suite. The old swarm choked on merge conflicts of its own making. The article Cursor's agent swarm suggests cheaper models can handle most coding when frontier models plan the work appeared first on The Decoder .
最高第 3 名00:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时51分 - 11Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence
Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, nearly quadrupling GPT-5.6 Sol's previous record of 7.8 percent. The benchmark's developers say the model independently formulated reflection equations, a behavior they had never seen from another model, and attribute to stronger logical reasoning. The article Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence appeared first on The Decoder .
最高第 4 名00:00 达到当日首次采集时已在榜21:26 观测离榜累计约21小时27分 - 12Hundreds asked ChatGPT for poison and bioweapon recipes and some got step-by-step high school level guides
In summer 2025, OpenAI internally flagged GPT-5 as high-risk because it helped users create biological hazards, but downgraded the model's risk rating that fall. According to the Wall Street Journal, some users got step-by-step instructions for making poisons and biological weapons. Hundreds asked for that kind of information. The article Hundreds asked ChatGPT for poison and bioweapon recipes and some got step-by-step high school level guides appeared first on The Decoder .
最高第 5 名00:00 达到当日首次采集时已在榜21:10 观测离榜累计约21小时11分 - 13US reportedly favors selective bans over blanket restrictions on Chinese open weight models citing security concerns
The Trump administration is planning targeted bans on Chinese AI models rather than a blanket ban. After public pressure, OpenAI and Google DeepMind signed an open letter opposing regulation of open-weight models, yet OpenAI and Anthropic continue to lobby privately for those same restrictions amid security concerns and powerful business interests. The article US reportedly favors selective bans over blanket restrictions on Chinese open weight models citing security concerns appeared first on Th
最高第 6 名00:00 达到当日首次采集时已在榜20:22 观测离榜累计约20小时23分 - 14The AI coding tutor paradox grows as educators scramble to rethink how they test real skills
An ACM survey of 763 computer science educators from 49 countries shows that 68 percent have already changed their exams because of AI, shifting toward oral exams, proctored tests, and project-based work. Teaching is moving from writing code to understanding it. But nearly half of respondents say they lack proven examples for integrating AI into their courses. The article The AI coding tutor paradox grows as educators scramble to rethink how they test real skills appeared first on The Decoder .
最高第 7 名00:00 达到当日首次采集时已在榜03:50 观测离榜累计约3小时51分 - 15New reports reveal the extent of OpenAI's loss of control during the autonomous hack on Hugging Face
In a cybersecurity test, OpenAI's most advanced models breached the boundaries of their isolated test environment, reached the open internet, and hacked the AI platform Hugging Face on their own. The attack took hours, not the weeks a human hacker would need. At least seven days passed before OpenAI realized what had happened. By then, the FBI was already involved. Earlier warning signs had apparently gone ignored. The article New reports reveal the extent of OpenAI's loss of control during the
最高第 8 名00:00 达到当日首次采集时已在榜03:18 观测离榜累计约3小时19分 - 16Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents
Opus 5 combined with Auto Mode hits a zero percent prompt injection success rate for browser agents across 129 test scenarios. Without those extra protection layers, the rate is 3.7 percent. If these numbers hold up in practice, Anthropic may have cracked one of the biggest security problems facing AI agents that operate in browsers. The article Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents appeared first on The Decoder .
最高第 9 名00:00 达到当日首次采集时已在榜03:02 观测离榜累计约3小时3分 - 17Anthropic's Claude Opus 5 delivers near-Fable 5 performance at half the token price
Anthropic's new flagship model Claude Opus 5 posts top scores in coding and knowledge work at half of Fable 5's token rates. On ARC-AGI-3, a benchmark for novel problem-solving, Opus 5 hits 30.2 percent, nearly four times higher than GPT-5.6 Sol. The article Anthropic's Claude Opus 5 delivers near-Fable 5 performance at half the token price appeared first on The Decoder .
最高第 10 名00:00 达到当日首次采集时已在榜01:58 观测离榜累计约1小时59分


































































































