全部/科技/实时热榜

The Decoder · 实时热榜

HISTORY2026年7月25日15 不同热搜
07/2308/21 有历史数据
DAILY UNIQUE TOPICS15 个热搜
  1. 01
    Sakana claims its AI model router Fugu Ultra v1.1 now beats Fable 5 without even including it in the pool

    Sakana AI has updated its Fugu Ultra AI router to version 1.1, claiming gains of up to 7.9 points over v1.0. Independent verification doesn't exist yet. The update adds a Claude Code-compatible endpoint. The service remains unavailable in the EU. The article Sakana claims its AI model router Fugu Ultra v1.1 now beats Fable 5 without even including it in the pool appeared first on The Decoder .

    最高第 100:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时54分
  2. 02
    Microsoft's open-weight AI push is so obviously an Azure play it hurts

    Microsoft, along with Meta, Nvidia, and more than 20 other companies, is pushing for open-weight AI models in an open letter. The strategic logic is simple: the more models running on Azure, the less Microsoft depends on expensive OpenAI and Anthropic models. The company is also swapping external models in products like Copilot for its in-house MAI family, which performs significantly worse in independent benchmarks. The article Microsoft's open-weight AI push is so obviously an Azure play it hu

    最高第 100:11 达到00:11 首次观测上榜当日结束时仍在榜累计约23小时42分
  3. 03
    Anthropic's Claude Opus 5 delivers near-Fable 5 performance at half the token price

    Anthropic's new flagship model Claude Opus 5 posts top scores in coding and knowledge work at half of Fable 5's token rates. On ARC-AGI-3, a benchmark for novel problem-solving, Opus 5 hits 30.2 percent, nearly four times higher than GPT-5.6 Sol. The article Anthropic's Claude Opus 5 delivers near-Fable 5 performance at half the token price appeared first on The Decoder .

    最高第 102:49 达到02:49 首次观测上榜当日结束时仍在榜累计约21小时4分
  4. 04
    Anthropic's Claude Opus 5 costs well below Fable 5 while matching or beating it across most benchmarks

    Anthropic's Claude Opus 5 leads the Artificial Analysis Intelligence Index with 61 points, edging out Claude Fable 5 and GPT-5.6 Sol. The model scores highest in analytical quality and coding, and costs up to half as much as Fable 5 at lower reasoning tiers. But the race at the top remains close. The article Anthropic's Claude Opus 5 costs well below Fable 5 while matching or beating it across most benchmarks appeared first on The Decoder .

    最高第 117:45 达到17:45 首次观测上榜当日结束时仍在榜累计约6小时8分
  5. 05
    Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents

    Opus 5 combined with Auto Mode hits a zero percent prompt injection success rate for browser agents across 129 test scenarios. Without those extra protection layers, the rate is 3.7 percent. If these numbers hold up in practice, Anthropic may have cracked one of the biggest security problems facing AI agents that operate in browsers. The article Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents appeared first on The Decoder .

    最高第 118:49 达到18:49 首次观测上榜当日结束时仍在榜累计约5小时4分
  6. 06
    New reports reveal the extent of OpenAI's loss of control during the autonomous hack on Hugging Face

    In a cybersecurity test, OpenAI's most advanced models breached the boundaries of their isolated test environment, reached the open internet, and hacked the AI platform Hugging Face on their own. The attack took hours, not the weeks a human hacker would need. At least seven days passed before OpenAI realized what had happened. By then, the FBI was already involved. Earlier warning signs had apparently gone ignored. The article New reports reveal the extent of OpenAI's loss of control during the

    最高第 122:01 达到22:01 首次观测上榜当日结束时仍在榜累计约1小时52分
  7. 07
    German AI consortium releases Soofi S, an open 30B model that tops benchmarks in both English and German

    The German consortium behind the AI model Soofi S has acknowledged in version 3.0 of its tech report that test questions from the science benchmark GPQA accidentally ended up in the training data. The community caught the error by examining the publicly available data. The team removed the benchmark from its evaluation and recalculated all results. The article German AI consortium releases Soofi S, an open 30B model that tops benchmarks in both English and German appeared first on The Decoder .

    最高第 200:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时54分
  8. 08
    Claude's voice mode now runs on Anthropic's most capable models across all platforms

    Voice conversations now run on the more powerful Opus and Sonnet models with access to Gmail, Google Calendar, and Slack. Claude is currently the only AI assistant that can compose and send emails directly by voice, giving it an edge over OpenAI and Google, whose voice output still sounds more natural. The article Claude's voice mode now runs on Anthropic's most capable models across all platforms appeared first on The Decoder .

    最高第 300:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时54分
  9. 09
    Kimi K3 trails frontier US models by a wide margin on cyber exploits, and distillation may explain why

    The British AI Security Institute and the U.S. Center for AI Standards and Innovation tested Moonshot AI's Kimi K3 on offensive cyber tasks. Kimi K3 scored 32 percent on ExploitBench, compared with 76 percent for leading U.S. models, while its safeguards failed to block exploit development or simulated attacks. The gap between its strong general benchmark scores and weaker cyber performance also fits allegations that Moonshot AI distilled Anthropic's models. The article Kimi K3 trails frontier U

    最高第 400:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时54分
  10. 10
    ChatGPT will give you worse health advice if you don't pay

    OpenAI is rolling out "Health in ChatGPT" to U.S. users, connecting Apple Health, medical records, and wellness apps. More than 300 million people already ask ChatGPT health questions every week, but paying users get better answers. The more powerful GPT-5.6 Sol model is reserved for premium subscribers, while free users are stuck with the weaker GPT-5.5 Instant. The article ChatGPT will give you worse health advice if you don't pay appeared first on The Decoder .

    最高第 500:00 达到当日首次采集时已在榜当日结束时仍在榜累计约23小时54分
  11. 11
    Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs

    Black Forest Labs has released Flux 3, a multimodal foundation model that learns from images, video, and audio and can generate video with native sound for the first time. BFL's own tests put it just ahead of market leader Seedance 2.0, though independent results aren't yet available. The company ultimately wants to build a world model and is already testing Flux 3 on robotics tasks. The article Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs appear

    最高第 600:00 达到当日首次采集时已在榜22:01 观测离榜累计约22小时2分
  12. 12
    One tampered ChatGPT link could spawn a rogue AI agent that took orders from an attacker every five minutes

    Zenity Labs uncovered "AgentForger," a vulnerability in OpenAI's Agent Builder that let a single manipulated ChatGPT link create an autonomous agent on an employee's behalf. The agent inherited the victim's identity and access rights, bypassed approval requirements through the malicious prompt, and pulled new instructions from the attacker's inbox every five minutes. The article One tampered ChatGPT link could spawn a rogue AI agent that took orders from an attacker every five minutes appeared f

    最高第 700:00 达到当日首次采集时已在榜18:49 观测离榜累计约18小时50分
  13. 13
    Poolside's Laguna S 2.1 is a small open-weight coding model that punches well above its size

    Poolside has released Laguna S 2.1, its third coding model in three months. Rather than rely on raw scale, the company trained it to keep checking its work, revise failed approaches, and avoid giving up too soon during long agentic sessions. The compact model beats several much larger rivals in benchmarks. Poolside says it also solved a math problem that had been open since 1975 for under 10 cents. The article Poolside's Laguna S 2.1 is a small open-weight coding model that punches well above it

    最高第 800:00 达到当日首次采集时已在榜17:45 观测离榜累计约17小时46分
  14. 14
    Google CEO Pichai says Gemini's next leap depends on building "much larger base models"

    Alphabet has raised its 2026 investment forecast to as much as $205 billion, saying demand continues to outpace spending. Google Cloud grew 82 percent in the second quarter. CEO Sundar Pichai says Google needs a larger base model for its next leap in AI and has kicked off an ambitious Gemini 4 training run. The article Google CEO Pichai says Gemini's next leap depends on building "much larger base models" appeared first on The Decoder .

    最高第 900:00 达到当日首次采集时已在榜02:49 观测离榜累计约2小时50分
  15. 15
    Anthropic's $1.5B piracy settlement with book authors is a record loss that hands AI labs their biggest legal win

    Anthropic has to pay $1.5 billion to book authors, the largest copyright settlement in class action history. But the payout is for downloading roughly 482,460 works from piracy databases, not for AI training itself. Judge Alsup had previously ruled that AI training on legally obtained books is "transformative" and falls under fair use. The settlement is actually a win for AI labs. The article Anthropic's $1.5B piracy settlement with book authors is a record loss that hands AI labs their biggest

    最高第 1000:00 达到当日首次采集时已在榜00:11 观测离榜累计约12分钟