全部/科技/实时热榜

The Decoder · 实时热榜

HISTORY2026年7月26日15 不同热搜
07/2308/21 有历史数据
DAILY UNIQUE TOPICS15 个热搜
  1. 01
    The AI coding tutor paradox grows as educators scramble to rethink how they test real skills

    An ACM survey of 763 computer science educators from 49 countries shows that 68 percent have already changed their exams because of AI, shifting toward oral exams, proctored tests, and project-based work. Teaching is moving from writing code to understanding it. But nearly half of respondents say they lack proven examples for integrating AI into their courses. The article The AI coding tutor paradox grows as educators scramble to rethink how they test real skills appeared first on The Decoder .

    最高第 115:13 达到15:13 首次观测上榜当日结束时仍在榜累计约8小时31分
  2. 02
    New reports reveal the extent of OpenAI's loss of control during the autonomous hack on Hugging Face

    In a cybersecurity test, OpenAI's most advanced models breached the boundaries of their isolated test environment, reached the open internet, and hacked the AI platform Hugging Face on their own. The attack took hours, not the weeks a human hacker would need. At least seven days passed before OpenAI realized what had happened. By then, the FBI was already involved. Earlier warning signs had apparently gone ignored. The article New reports reveal the extent of OpenAI's loss of control during the

    最高第 100:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时45分
  3. 03
    US reportedly favors selective bans over blanket restrictions on Chinese open weight models citing security concerns

    The Trump administration is planning targeted bans on Chinese AI models rather than a blanket ban. After public pressure, OpenAI and Google DeepMind signed an open letter opposing regulation of open-weight models, yet OpenAI and Anthropic continue to lobby privately for those same restrictions amid security concerns and powerful business interests. The article US reportedly favors selective bans over blanket restrictions on Chinese open weight models citing security concerns appeared first on Th

    最高第 116:01 达到16:01 首次观测上榜当日结束时仍在榜累计约7小时43分
  4. 04
    Hundreds asked ChatGPT for poison and bioweapon recipes and some got step-by-step high school level guides

    In summer 2025, OpenAI internally flagged GPT-5 as high-risk because it helped users create biological hazards, but downgraded the model's risk rating that fall. According to the Wall Street Journal, some users got step-by-step instructions for making poisons and biological weapons. Hundreds asked for that kind of information. The article Hundreds asked ChatGPT for poison and bioweapon recipes and some got step-by-step high school level guides appeared first on The Decoder .

    最高第 116:49 达到16:49 首次观测上榜当日结束时仍在榜累计约6小时55分
  5. 05
    Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence

    Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, nearly quadrupling GPT-5.6 Sol's previous record of 7.8 percent. The benchmark's developers say the model independently formulated reflection equations, a behavior they had never seen from another model, and attribute to stronger logical reasoning. The article Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence appeared first on The Decoder .

    最高第 117:53 达到17:53 首次观测上榜当日结束时仍在榜累计约5小时51分
  6. 06
    Cursor's agent swarm suggests cheaper models can handle most coding when frontier models plan the work

    Cursor asked its upgraded agent swarm and its predecessor to rebuild SQLite in Rust using only the documentation, with no source code or internet access. Every configuration of the new system, which separates planners from workers, eventually scored 100 percent on the test suite. The old swarm choked on merge conflicts of its own making. The article Cursor's agent swarm suggests cheaper models can handle most coding when frontier models plan the work appeared first on The Decoder .

    最高第 123:13 达到23:13 首次观测上榜当日结束时仍在榜累计约32分钟
  7. 07
    Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents

    Opus 5 combined with Auto Mode hits a zero percent prompt injection success rate for browser agents across 129 test scenarios. Without those extra protection layers, the rate is 3.7 percent. If these numbers hold up in practice, Anthropic may have cracked one of the biggest security problems facing AI agents that operate in browsers. The article Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents appeared first on The Decoder .

    最高第 200:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时45分
  8. 08
    Anthropic's Claude Opus 5 delivers near-Fable 5 performance at half the token price

    Anthropic's new flagship model Claude Opus 5 posts top scores in coding and knowledge work at half of Fable 5's token rates. On ARC-AGI-3, a benchmark for novel problem-solving, Opus 5 hits 30.2 percent, nearly four times higher than GPT-5.6 Sol. The article Anthropic's Claude Opus 5 delivers near-Fable 5 performance at half the token price appeared first on The Decoder .

    最高第 300:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时45分
  9. 09
    Anthropic's Claude Opus 5 costs well below Fable 5 while matching or beating it across most benchmarks

    Anthropic's Claude Opus 5 leads the Artificial Analysis Intelligence Index with 61 points, edging out Claude Fable 5 and GPT-5.6 Sol. The model scores highest in analytical quality and coding, and costs up to half as much as Fable 5 at lower reasoning tiers. But the race at the top remains close. The article Anthropic's Claude Opus 5 costs well below Fable 5 while matching or beating it across most benchmarks appeared first on The Decoder .

    最高第 400:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时45分
  10. 10
    Microsoft's open-weight AI push is so obviously an Azure play it hurts

    Microsoft, along with Meta, Nvidia, and more than 20 other companies, is pushing for open-weight AI models in an open letter. The strategic logic is simple: the more models running on Azure, the less Microsoft depends on expensive OpenAI and Anthropic models. The company is also swapping external models in products like Copilot for its in-house MAI family, which performs significantly worse in independent benchmarks. The article Microsoft's open-weight AI push is so obviously an Azure play it hu

    最高第 500:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时45分
  11. 11
    Sakana claims its AI model router Fugu Ultra v1.1 now beats Fable 5 without even including it in the pool

    Sakana AI has updated its Fugu Ultra AI router to version 1.1, claiming gains of up to 7.9 points over v1.0. Independent verification doesn't exist yet. The update adds a Claude Code-compatible endpoint. The service remains unavailable in the EU. The article Sakana claims its AI model router Fugu Ultra v1.1 now beats Fable 5 without even including it in the pool appeared first on The Decoder .

    最高第 600:00 达到当日首次采集时已在榜23:13 观测离榜累计约23小时13分
  12. 12
    German AI consortium releases Soofi S, an open 30B model that tops benchmarks in both English and German

    The German consortium behind the AI model Soofi S has acknowledged in version 3.0 of its tech report that test questions from the science benchmark GPQA accidentally ended up in the training data. The community caught the error by examining the publicly available data. The team removed the benchmark from its evaluation and recalculated all results. The article German AI consortium releases Soofi S, an open 30B model that tops benchmarks in both English and German appeared first on The Decoder .

    最高第 700:00 达到当日首次采集时已在榜17:53 观测离榜累计约17小时54分
  13. 13
    Claude's voice mode now runs on Anthropic's most capable models across all platforms

    Voice conversations now run on the more powerful Opus and Sonnet models with access to Gmail, Google Calendar, and Slack. Claude is currently the only AI assistant that can compose and send emails directly by voice, giving it an edge over OpenAI and Google, whose voice output still sounds more natural. The article Claude's voice mode now runs on Anthropic's most capable models across all platforms appeared first on The Decoder .

    最高第 800:00 达到当日首次采集时已在榜16:49 观测离榜累计约16小时50分
  14. 14
    Kimi K3 trails frontier US models by a wide margin on cyber exploits, and distillation may explain why

    The British AI Security Institute and the U.S. Center for AI Standards and Innovation tested Moonshot AI's Kimi K3 on offensive cyber tasks. Kimi K3 scored 32 percent on ExploitBench, compared with 76 percent for leading U.S. models, while its safeguards failed to block exploit development or simulated attacks. The gap between its strong general benchmark scores and weaker cyber performance also fits allegations that Moonshot AI distilled Anthropic's models. The article Kimi K3 trails frontier U

    最高第 900:00 达到当日首次采集时已在榜16:01 观测离榜累计约16小时2分
  15. 15
    ChatGPT will give you worse health advice if you don't pay

    OpenAI is rolling out "Health in ChatGPT" to U.S. users, connecting Apple Health, medical records, and wellness apps. More than 300 million people already ask ChatGPT health questions every week, but paying users get better answers. The more powerful GPT-5.6 Sol model is reserved for premium subscribers, while free users are stuck with the weaker GPT-5.5 Instant. The article ChatGPT will give you worse health advice if you don't pay appeared first on The Decoder .

    最高第 1000:00 达到当日首次采集时已在榜15:13 观测离榜累计约15小时14分