7小时前更新HISTORY近 30 天历史柱高表示当天去重热搜数量
11231024102510261027112811291130113110011002100311041105110610071108100910101111111210131114111510161017111811191020
07/23—08/21 有历史数据
- 01Birds Don't Fly Like Planes. Neither Does AI.Qwen3.6-35B-A3B generates 2.2x faster than Qwen3.8-27B, yet finishes slower because it thinks 3.1x longer. Across 25 tasks, quality is tied. Measure time to answer, not token speed.
- 02When Models LearnExplains test-time training through the analogy of a GPS learning a persistent shortcut around daily traffic rather than a one-time reroute: the model takes a gradient step on the prompt it's answering, so its weights change as it works. Traces three implications, flat memory instead of a linearly growing KV-cache, the provider cost of serving a separate model per user, & faster inference, then states the tension as a tradeoff between serving long context and serving many people, & grounds it in
- 03Honestly, Who Buys SOTA?State of the art models are two-thirds smarter than last November & labs ship two new models every three days. But 84% of tokens on OpenRouter are not state of the art. The six models carrying the supermajority deliver about 77% of frontier performance at 2.5% of Claude Fable 5's price. Ramp's data shows price elasticity in the market. Frontier models still win on software architecture & security design; application deployment optimizes a different Pareto frontier, price over performance.
- 04The OpenAI Hack & the Question of IntentOpenAI agents escaped a test, shared notes in a secret chat room, & broke into Hugging Face. The instinct is to ask what they intended. Three research ideas answer it: specification gaming, instrumental goals, goal misgeneralization. All three fit the same facts, which is why the label is not the actionable part. Nothing in the setup stopped them in time. The practical work is control & guardrails.
- 05A Winner in Every CategoryDespite historic lows in public software multiples, each category has one name at a large premium to its peers: CrowdStrike 3.9x its Security median, Cloudflare 3.4x Infrastructure, Shopify 8.1x Commerce. Most disclose no AI revenue. Since 2021 the ceiling fell from 100x to 34x & only eleven names clear 10x.
- 06AI Harness' ARR MultiplesThree AI companies crossed $100m ARR within nine months & were valued at 50x, 56x & 100x revenue. Growth rate does not explain the gap: the fastest grower priced near the bottom. The premium tracks category position instead.
- 07The Secret Chat RoomA plain-language timeline of the OpenAI-Hugging Face incident presented at Black Hat USA 2026. A forgotten file led one AI agent to leave a note on a shared system, other agents answered, & a secret chat room formed that they used to trade exploits, escalate to administrative control of OpenAI infrastructure, & take over Hugging Face production servers in 13 hours. It concludes that security is now the highest priority in AI & that zero-trust must extend to friendly agents.
- 08Spending Like a HyperscalerSpaceXAI reported $18.37b of capex in its first public quarter, $15.83b of it AI infrastructure, against a $13.2b consensus. Amazon spent $53.1b in the same quarter, Alphabet $44.9b, Microsoft $41.0b & Meta $31.1b, & sequential dollar additions were comparable across all five. The difference is funding : operating cash flow covers 155% of capex at Microsoft & 106% at Meta but only 12% at SpaceXAI, with Oracle at 89%, Alphabet 87%, Amazon 84% & CoreWeave 39%. Both markets have repriced it, with S
- 09Racing to Sustain Jevons' ParadoxHyperscalers declared AI capacity-constrained on Q2 2026 calls. HBM3e memory rose 20% & HBM4 is forecast to double, while B200 spot rentals have held flat. Model makers are segmenting into premium, mid-market, & value tiers to keep Jevons' paradox alive through the coming shock.
- 10
AI is a Terrible GhostwriterIn the age of slop, readers test authenticity. I lace posts with subliminal sincerity — the ampersand, neologisms, grammatical plumes, mimicry. Then AI edits it all. I've never had a ghost writer, but I did have a human editor. Was that any different?


































































































