全部/科技/实时热榜

The Decoder · 实时热榜

HISTORY2026年7月27日12 不同热搜
07/2308/21 有历史数据
DAILY UNIQUE TOPICS12 个热搜
  1. 01
    Cursor's agent swarm suggests cheaper models can handle most coding when frontier models plan the work

    Cursor asked its upgraded agent swarm and its predecessor to rebuild SQLite in Rust using only the documentation, with no source code or internet access. Every configuration of the new system, which separates planners from workers, eventually scored 100 percent on the test suite. The old swarm choked on merge conflicts of its own making. The article Cursor's agent swarm suggests cheaper models can handle most coding when frontier models plan the work appeared first on The Decoder .

    最高第 100:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时45分
  2. 02
    Shared Claude chats were reportedly showing up in search engines

    Shared conversations with Anthropic's Claude chatbot briefly appeared in Google search results because the pages lacked a noindex tag. Users said some chats contained crypto keys and legal questions. OpenAI made the same mistake last year. The article Shared Claude chats were reportedly showing up in search engines appeared first on The Decoder .

    最高第 116:16 达到16:16 首次观测上榜当日结束时仍在榜累计约7小时28分
  3. 03
    METR introduces a new metric to calculate exactly when AI agents become more expensive than humans

    METR's new metric, the "expenditure horizon," puts a dollar figure on how cost-effective AI agents are at solving problems. Early results on the NanoGPT speedrun are underwhelming, the metric has blind spots, and the newest generation of models could change the picture. The article METR introduces a new metric to calculate exactly when AI agents become more expensive than humans appeared first on The Decoder .

    最高第 120:32 达到20:32 首次观测上榜当日结束时仍在榜累计约3小时12分
  4. 04
    Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence

    Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, nearly quadrupling GPT-5.6 Sol's previous record of 7.8 percent. The benchmark's developers say the model independently formulated reflection equations, a behavior they had never seen from another model, and attribute to stronger logical reasoning. The article Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence appeared first on The Decoder .

    最高第 200:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时45分
  5. 05
    Hundreds asked ChatGPT for poison and bioweapon recipes and some got step-by-step high school level guides

    In summer 2025, OpenAI internally flagged GPT-5 as high-risk because it helped users create biological hazards, but downgraded the model's risk rating that fall. According to the Wall Street Journal, some users got step-by-step instructions for making poisons and biological weapons. Hundreds asked for that kind of information. The article Hundreds asked ChatGPT for poison and bioweapon recipes and some got step-by-step high school level guides appeared first on The Decoder .

    最高第 300:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时45分
  6. 06
    US reportedly favors selective bans over blanket restrictions on Chinese open weight models citing security concerns

    The Trump administration is planning targeted bans on Chinese AI models rather than a blanket ban. After public pressure, OpenAI and Google DeepMind signed an open letter opposing regulation of open-weight models, yet OpenAI and Anthropic continue to lobby privately for those same restrictions amid security concerns and powerful business interests. The article US reportedly favors selective bans over blanket restrictions on Chinese open weight models citing security concerns appeared first on Th

    最高第 400:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时45分
  7. 07
    The AI coding tutor paradox grows as educators scramble to rethink how they test real skills

    An ACM survey of 763 computer science educators from 49 countries shows that 68 percent have already changed their exams because of AI, shifting toward oral exams, proctored tests, and project-based work. Teaching is moving from writing code to understanding it. But nearly half of respondents say they lack proven examples for integrating AI into their courses. The article The AI coding tutor paradox grows as educators scramble to rethink how they test real skills appeared first on The Decoder .

    最高第 500:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时45分
  8. 08
    New reports reveal the extent of OpenAI's loss of control during the autonomous hack on Hugging Face

    In a cybersecurity test, OpenAI's most advanced models breached the boundaries of their isolated test environment, reached the open internet, and hacked the AI platform Hugging Face on their own. The attack took hours, not the weeks a human hacker would need. At least seven days passed before OpenAI realized what had happened. By then, the FBI was already involved. Earlier warning signs had apparently gone ignored. The article New reports reveal the extent of OpenAI's loss of control during the

    最高第 600:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时45分
  9. 09
    Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents

    Opus 5 combined with Auto Mode hits a zero percent prompt injection success rate for browser agents across 129 test scenarios. Without those extra protection layers, the rate is 3.7 percent. If these numbers hold up in practice, Anthropic may have cracked one of the biggest security problems facing AI agents that operate in browsers. The article Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents appeared first on The Decoder .

    最高第 700:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时45分
  10. 10
    Anthropic's Claude Opus 5 delivers near-Fable 5 performance at half the token price

    Anthropic's new flagship model Claude Opus 5 posts top scores in coding and knowledge work at half of Fable 5's token rates. On ARC-AGI-3, a benchmark for novel problem-solving, Opus 5 hits 30.2 percent, nearly four times higher than GPT-5.6 Sol. The article Anthropic's Claude Opus 5 delivers near-Fable 5 performance at half the token price appeared first on The Decoder .

    最高第 800:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时45分
  11. 11
    Anthropic's Claude Opus 5 costs well below Fable 5 while matching or beating it across most benchmarks

    Anthropic's Claude Opus 5 leads the Artificial Analysis Intelligence Index with 61 points, edging out Claude Fable 5 and GPT-5.6 Sol. The model scores highest in analytical quality and coding, and costs up to half as much as Fable 5 at lower reasoning tiers. But the race at the top remains close. The article Anthropic's Claude Opus 5 costs well below Fable 5 while matching or beating it across most benchmarks appeared first on The Decoder .

    最高第 900:00 达到当日首次采集时已在榜20:32 观测离榜累计约20小时33分
  12. 12
    Microsoft's open-weight AI push is so obviously an Azure play it hurts

    Microsoft, along with Meta, Nvidia, and more than 20 other companies, is pushing for open-weight AI models in an open letter. The strategic logic is simple: the more models running on Azure, the less Microsoft depends on expensive OpenAI and Anthropic models. The company is also swapping external models in products like Copilot for its in-house MAI family, which performs significantly worse in independent benchmarks. The article Microsoft's open-weight AI push is so obviously an Azure play it hu

    最高第 1000:00 达到当日首次采集时已在榜16:16 观测离榜累计约16小时17分