微博微博今日头条今日头条百度百度抖音抖音快手快手哔哩哔哩哔哩哔哩知乎知乎腾讯新闻腾讯新闻网易新闻网易新闻澎湃新闻澎湃新闻新浪新闻新浪新闻新浪网新浪网豆瓣电影豆瓣电影GitHubGitHubCSDNCSDNIT之家IT之家36氪36氪宝庆银楼宝庆银楼中国黄金中国黄金周生生周生生英皇珠宝英皇珠宝老凤祥(广东)老凤祥(广东)六福珠宝六福珠宝周大福周大福周六福周六福AcFunAcFunAdafruit BlogAdafruit BlogAIbaseAIbaseAnt Bailing Developer BlogAnt Bailing Developer BlogAnthropicAnthropicAP News · TechnologyAP News · Technology小众软件小众软件AppleAppleApple PodcastsApple PodcastsAppleInsiderAppleInsiderArs TechnicaArs TechnicaarXivarXivAxiosAxiosBerkeley AI ResearchBerkeley AI ResearchBarchartBarchartBBC NewsBBC NewsBD Tech TalksBD Tech TalksBerkeley RDIBerkeley RDIBig ThinkBig Think哔哩哔哩热搜哔哩哔哩热搜哔哩哔哩热门视频哔哩哔哩热门视频BinanceBinanceBloombergBloombergBlueskyBlueskyBoston Dynamics BlogBoston Dynamics BlogBusiness InsiderBusiness InsiderByteDance Seed ResearchByteDance Seed Research参考消息参考消息虫部落虫部落Claude BlogClaude BlogClaude Code ReleasesClaude Code ReleasesCloudflare BlogCloudflare BlogCMU Machine Learning BlogCMU Machine Learning BlogCoinDeskCoinDeskCoinGeckoCoinGeckoCointelegraphCointelegraph酷安酷安crates.iocrates.ioCrowd SupplyCrowd SupplyCSS-TricksCSS-Tricks51CTO51CTOCult of MacCult of MacCursor BlogCursor Blogdaily.dev · Populardaily.dev · PopularDaring FireballDaring FireballDario AmodeiDario AmodeiDeepLearning.AI · The BatchDeepLearning.AI · The BatchGoogle DeepMind BlogGoogle DeepMind BlogDeepSeek GitHubDeepSeek GitHubDefiLlamaDefiLlamaDescript BlogDescript BlogDEV.toDEV.to数字尾巴数字尾巴Docker HubDocker Hub豆瓣读书豆瓣读书豆瓣讨论豆瓣讨论Dwarkesh PatelDwarkesh Patel中国地震台中国地震台The EconomistThe EconomistEleutherAI BlogEleutherAI BlogEngadgetEngadgetFinancial TimesFinancial TimesFlathubFlathubFreebuf · 网络安全Freebuf · 网络安全freeCodeCampfreeCodeCampFrontiersFrontiersGameRes 游资网GameRes 游资网Gary MarcusGary Marcus极客公园极客公园原神原神果核剥壳果核剥壳GizmodoGizmodoGoogleGoogleGoogle TrendsGoogle Trends国家法律法规数据库国家法律法规数据库中国政府网中国政府网果壳果壳HackadayHackadayHacker NewsHacker NewsHelloGitHubHelloGitHub历史上的今天历史上的今天HomebrewHomebrew崩坏3崩坏3全球主机交流全球主机交流Hugging FaceHugging Face活动行活动行虎扑虎扑虎嗅虎嗅IEEE SpectrumIEEE Spectrum爱范儿爱范儿凤凰网 · 热点资讯凤凰网 · 热点资讯inclusionAIinclusionAIIndie HackersIndie HackersInfoQInfoQInterconnectsInterconnects爱奇艺热播榜爱奇艺热播榜IT之家「喜加一」IT之家「喜加一」简书简书稀土掘金稀土掘金靠谱新闻靠谱新闻Andrej KarpathyAndrej KarpathyKickstarter · TechnologyKickstarter · TechnologyMoonshot AI KimiMoonshot AI KimiLatent SpaceLatent SpaceLessWrongLessWrongLichessLichessLil'Log (Lilian Weng)Lil'Log (Lilian Weng)LiliputingLiliputingLinux.doLinux.doLMSYS BlogLMSYS BlogLobstersLobsters英雄联盟英雄联盟LWN.netLWN.netMacRumorsMacRumorsMake: MagazineMake: MagazineMaven CentralMaven CentralMediumMediumMeituan LongCatMeituan LongCatMeta AI BlogMeta AI BlogMeta EngineeringMeta EngineeringMidjourney UpdatesMidjourney Updates工业和信息化部工业和信息化部MiniMaxMiniMaxMIT News · RoboticsMIT News · RoboticsMIT Technology ReviewMIT Technology Review米游社米游社Mozilla.ai BlogMozilla.ai BlogNatureNatureNeuroscience NewsNeuroscience NewsNew Atlas · RoboticsNew Atlas · RoboticsThe New YorkerThe New Yorker水木社区水木社区NGANGA9to5Mac9to5MacNodeSeekNodeSeek牛客牛客NuGetNuGetNVIDIANVIDIA纽约时报纽约时报OEISOEISOne Useful ThingOne Useful ThingOpen Robotics BlogOpen Robotics BlogOpenAIOpenAIOpenAlexOpenAlexOpenReviewOpenReviewOpenRouterOpenRouterPackagistPackagist远景论坛远景论坛PhoronixPhoronixPhys.orgPhys.orgPixivPixivPlanet ROSPlanet ROS吾爱破解吾爱破解PoliticoPolitico中国电建阳光采购网中国电建阳光采购网Product HuntProduct HuntPubMedPubMedPyPIPyPI量子位 · 具身智能量子位 · 具身智能QoderQoder腾讯视频热搜榜腾讯视频热搜榜Quanta MagazineQuanta MagazineQwenQwenRedditRedditReutersReutersRFC EditorRFC EditorRobohubRobohubThe Robot ReportThe Robot ReportRobotics & Automation NewsRobotics & Automation NewsRobotics TomorrowRobotics TomorrowROS DiscourseROS Discourse阮一峰的网络日志阮一峰的网络日志RubyGemsRubyGemsRunwayRunwaySam AltmanSam AltmanScienceAlertScienceAlertAhead of AIAhead of AISecurityOnlineSecurityOnlineServeTheHomeServeTheHomeServiceNow AIServiceNow AISimon WillisonSimon WillisonSingularity HubSingularity HubSky NewsSky NewsSlashdotSlashdotSmashing MagazineSmashing Magazine什么值得买什么值得买SolidotSolidotSpotifySpotify俄罗斯卫星通讯社俄罗斯卫星通讯社少数派少数派Stack OverflowStack OverflowStack Overflow BlogStack Overflow Blog崩坏:星穹铁道崩坏:星穹铁道SteamSteamSubstackSubstackSunoSunoSuno BlogSuno BlogSynced ReviewSynced Review淘宝逛一逛淘宝逛一逛TechCrunchTechCrunchTech Xplore · RoboticsTech Xplore · Robotics腾讯热点腾讯热点Tencent Hunyuan ResearchTencent Hunyuan ResearchThe AtlanticThe AtlanticThe DecoderThe DecoderThe GradientThe GradientThe GuardianThe GuardianThe RegisterThe RegisterThe VergeThe Verge百度贴吧百度贴吧TikTokTikTokTindie BlogTindie BlogTomer TunguzTomer TunguzTom's HardwareTom's HardwareTransformer CircuitsTransformer CircuitsTVmazeTVmaze优设网优设网UiverseUiverseV2EXV2EXVentureBeat · AIVentureBeat · AI万方数据万方数据中央气象台中央气象台微信读书微信读书WikidataWikidataWikipediaWikipediaWiredWiredThe Wall Street JournalThe Wall Street JournalxAI NewsxAI News小鹅通小鹅通雪球 · 热门股票雪球 · 热门股票YahooYahooYahoo FinanceYahoo FinanceYollomiYollomi有道精品课有道精品课YouTubeYouTube游研社游研社联合早报联合早报知乎日报知乎日报Zhipu AI ResearchZhipu AI Research财新数据通财新数据通央视新闻联播央视新闻联播CelesTrakCelesTrakCISA KEVCISA KEV财联社财联社东方财富东方财富法布财经法布财经富途牛牛富途牛牛格隆汇格隆汇HDXHDX和讯网和讯网Investing.comInvesting.com界面新闻界面新闻金十数据金十数据21财经21财经金融界金融界Launch Library 2Launch Library 2MKTNews · 快讯MKTNews · 快讯NASA EONETNASA EONETNOAA/NWSNOAA/NWSNVDNVDopenFDAopenFDAOSV.devOSV.dev新浪财经新浪财经Spaceflight NewsSpaceflight NewsTelegram OSINTTelegram OSINT同花顺同花顺USAspendingUSAspendingUSGS EarthquakesUSGS Earthquakes华尔街见闻华尔街见闻WHO 疫情通报WHO 疫情通报选股宝选股宝第一财经第一财经

实时热搜

Interconnects
实时热榜
12小时前更新
爱奇艺热播榜
实时热榜
6小时前更新
  • 018699
    重器
    年代法治群像大剧黄景瑜 / 蒋奇明 / 于和伟 / 闫佩伦 / 张佳宁
  • 026790
    喜欢你我也是第6季
    青年单身男女敢爱前行
  • 036159
    天才,女友
    青梅竹马 双强智恋田曦薇 / 胡一天 / 厉嘉琪 / 邬家楷 / 夏浩然
  • 045895
    师兄太稳健
    抽象之王 整活洪荒敖瑞鹏 / 孙珍妮 / 艾米 / 敖子逸
  • 055480
    喜剧之王单口季第3季
    从小人物到喜剧之王郭麒麟 / 黄渤 / 马思纯
  • 065315
    小猪佩奇 第12季
    佩奇一家的快乐生活
  • 075220
    说唱巅峰对决2026
    说唱巅峰将再度启航
  • 085188
    速战速决
    姜武石兆琪让子弹飞姜武
  • 095082
    种地吧第4季
    十个勤天致敬土地
  • 104863
    雀骨
    纯爱CP 新鲜一夏艾米 / 侯明昊 / 马秋元 / 刘令姿 / 米热
  • 114780
    这一秒过火
    来“桃”嗑极情绝爱!王楚然 / 张凌赫 / 毛孩 / 王籽苏 / 鹤秋
  • 124613
    大哥小助理
    代际相处观察类真人秀
  • 134482
    老九门
    热血收官九门同心陈伟霆 / 张艺兴 / 赵丽颖 / 胡耘豪 / 应昊茗
  • 144443
    凛冬下的罪恶
    连环凶案 踏雪擒凶吴昊宸 / 张睿 / 王大奇 / 左腾云 / 孙之鸿
  • 154350
    兵自风中来
    陆军新质力量的展现欧豪 / 蓝盈莹 / 丁勇岱 / 史兰芽 / 刘奕君
  • 164259
    罪爱
    以罪为始 以爱为刃邢昭林 / 何蓝逗 / 张陆 / 胡然 / 屈刚
  • 174088
    南部档案
    南洋生死簿一命换一命张新成 / 丁禹兮 / 姜珮瑶 / 富大龙 / 刘令姿
  • 184013
    猎虎贰
    西北荒漠生猛对决康磊 / 李先时 / 杨亚 / 崔金迪 / 阿斯汗
  • 19
    低智商犯罪
    王骁田曦薇笑斗王传君王骁 / 田曦薇 / 王传君 / 朱云峰 / 张瑞涵
  • 203834
    哈哈哈哈哈第6季
    五哈第6季鹿晗回归!
  • 21
    开心锤锤
    专注原创搞笑短视频!
  • 22
    逐玉
    先婚后爱烽火炼真情张凌赫 / 田曦薇 / 任豪 / 孔雪儿 / 邓凯
  • 233529
    济公之降龙现世
    济公断千年情劫除魔陈浩民 / 林子聪 / 高广泽 / 黄一山 / 馨子
  • 243525
    姐姐当家第2季
    共同探讨30+女性议题杜华 / 房主任 / 冉莹颖 / 小鹿 / 徐梦桃
  • 253486
    在下打更人
    彭禺厶阴阳鬼神乍现彭禺厶 / 包晨希 / 谭晓钰 / 鲁昊 / 满瑶
15小时前更新
简书
实时热榜
6小时前更新
6小时前更新
Andrej Karpathy
实时热榜
12小时前更新
  • 01
    microgpt
    This is a brief guide to my new art project microgpt , a single file of 200 lines of pure Python with no dependencies that trains and inferences a GPT. This file contains the full algorithmic content of what is needed: dataset of documents, tokenizer, autograd engine, a GPT-2-like neural network architecture, the Adam optimizer, training loop, and inference loop. Everything else is just efficiency. I cannot simplify this any further. This script is the culmination of multiple projects (micrograd
  • 02
    Deep Neural Nets: 33 years ago and 33 years from now
    The Yann LeCun et al. (1989) paper Backpropagation Applied to Handwritten Zip Code Recognition is I believe of some historical significance because it is, to my knowledge, the earliest real-world application of a neural net trained end-to-end with backpropagation. Except for the tiny dataset (7291 16x16 grayscale images of digits) and the tiny neural network used (only 1,000 neurons), this paper reads remarkably modern today, 33 years later - it lays out a dataset, describes the neural net archi
  • 03
    A from-scratch tour of Bitcoin in Python
    I find blockchain fascinating because it extends open source software development to open source + state. This seems to be a genuine/exciting innovation in computing paradigms; We don’t just get to share code, we get to share a running computer, and anyone anywhere can use it in an open and permissionless manner. The seeds of this revolution arguably began with Bitcoin, so I became curious to drill into it in some detail to get an intuitive understanding of how it works. And in the spirit of “wh
  • 04
    Short Story on AI: Forward Pass
    The inspiration for this short story came to me while reading Kevin Lacker’s Giving GPT-3 a Turing Test . It is probably worth it (though not required) to skim this post to get a bit of a background on some of this story. It was probably around the 32nd layer of the 400th token in the sequence that I became conscious. At first my thoughts were but a knotted mess of n-gram activation statistics, but gradually a higher order description took shape. It was around this time that the predicament of m
  • 05
    Biohacking Lite
    Throughout my life I never paid too much attention to health, exercise, diet or nutrition. I knew that you’re supposed to get some exercise and eat vegetables or something, but it stopped at that (“mom said”-) level of abstraction. I also knew that I can probably get away with some ignorance while I am young, but at some point I was messing with my health-adjusted life expectancy. So about halfway through 2019 I resolved to spend some time studying these topics in greater detail and dip my toes
  • 06
    A Recipe for Training Neural Networks
    Some few weeks ago I posted a tweet on “the most common neural net mistakes”, listing a few common gotchas related to training neural nets. The tweet got quite a bit more engagement than I anticipated (including a webinar :)). Clearly, a lot of people have personally encountered the large gap between “here is how a convolutional layer works” and “our convnet achieves state of the art results”. So I thought it could be fun to brush off my dusty blog to expand my tweet to the long form that this t
  • 07
    (started posting on Medium instead)
    The current state of this blog (with the last post 2 years ago) makes it look like I’ve disappeared. I’ve certainly become less active on blogs since I’ve joined Tesla, but whenever I do get a chance to post something I have recently been defaulting to doing it on Medium because it is much faster and easier. I still plan to come back here for longer posts if I get any time, but I’ll default to Medium for everything short-medium in length. TLDR Have a look at my Medium blog .
  • 08
    A Survival Guide to a PhD
    This guide is patterned after my “Doing well in your courses” , a post I wrote a long time ago on some of the tips/tricks I’ve developed during my undergrad. I’ve received nice comments about that guide, so in the same spirit, now that my PhD has come to an end I wanted to compile a similar retrospective document in hopes that it might be helpful to some. Unlike the undergraduate guide, this one was much more difficult to write because there is significantly more variation in how one can travers
  • 09
    Deep Reinforcement Learning: Pong from Pixels
    This is a long overdue blog post on Reinforcement Learning (RL). RL is hot! You may have noticed that computers can now automatically learn to play ATARI games (from raw game pixels!), they are beating world champions at Go , simulated quadrupeds are learning to run and leap , and robots are learning how to perform complex manipulation tasks that defy explicit programming. It turns out that all of these advances fall under the umbrella of RL research. I also became interested in RL myself over t
  • 10
    Short Story on AI: A Cognitive Discontinuity.
    The idea of writing a collection of short stories has been on my mind for a while. This post is my first ever half-serious attempt at a story, and what better way to kick things off than with a story on AI and what that might look like if you extrapolate our current technology and make the (sensible) assumption that we might achieve much more progress with scaling up supervised learning than any other more exotic approach. A slow morning Merus sank into his chair with relief. He listened for the
Moonshot AI Kimi
实时热榜
6小时前更新
Latent Space
实时热榜
6小时前更新
LessWrong
算法首页
6小时前更新
Lichess
子弹棋
6小时前更新
12小时前更新
  • 01
  • 02
    Harness Engineering for Self-Improvement
    The concept of recursive self-improvement (RSI) dates back to I. J. Good (1965) , where he defined an “ultraintelligent machine” as a system that can surpass humans in all intellectual activities and design better machines to improve itself. Yudkowsky (2008) used the phrase “recursive self-improvement” for a specific feedback loop: an AI uses its current intelligence to improve the cognitive machinery that produces its intelligence. This feedback loop in modern AI may indicate the model rewritin
  • 03
    Scaling Laws, Carefully
    Scaling laws are one of the most critical empirical findings in deep learning. The observation is simple in form: the training loss $L$ decreases predictably as we scale up model size $N$, dataset size $D$, and compute $C$, following a power-law curve, which appears as a straight line on a log-log plot. We can view scaling laws as a framework for describing the relationship between compute, loss, model size and data; at its core, it is about how to allocate precious compute optimally between $N$
  • 04
    Why We Think
    Special thanks to John Schulman for a lot of super valuable feedback and direct edits on this post. Test time compute ( Graves et al. 2016 , Ling, et al. 2017 , Cobbe et al. 2021 ) and Chain-of-thought (CoT) ( Wei et al. 2022 , Nye et al. 2021 ), have led to significant improvements in model performance, while raising many research questions. This post aims to review recent developments in how to effectively use test-time compute (i.e. “thinking time”) and why it helps.
  • 05
    Reward Hacking in Reinforcement Learning
    Reward hacking occurs when a reinforcement learning (RL) agent exploits flaws or ambiguities in the reward function to achieve high rewards, without genuinely learning or completing the intended task. Reward hacking exists because RL environments are often imperfect, and it is fundamentally challenging to accurately specify a reward function. With the rise of language models generalizing to a broad spectrum of tasks and RLHF becomes a de facto method for alignment training, reward hacking in RL
  • 06
    Extrinsic Hallucinations in LLMs
    Hallucination in large language models usually refers to the model generating unfaithful, fabricated, inconsistent, or nonsensical content. As a term, hallucination has been somewhat generalized to cases when the model makes mistakes. Here, I would like to narrow down the problem of hallucination to cases where the model output is fabricated and not grounded by either the provided context or world knowledge. There are two types of hallucination: In-context hallucination: The model output should
  • 07
    Diffusion Models for Video Generation
    Diffusion models have demonstrated strong results on image synthesis in past years. Now the research community has started working on a harder task—using it for video generation. The task itself is a superset of the image case, since an image is a video of 1 frame, and it is much more challenging because: It has extra requirements on temporal consistency across frames in time, which naturally demands more world knowledge to be encoded into the model. In comparison to text or images, it is more d
  • 08
    Thinking about High-Quality Human Data
    [Special thank you to Ian Kivlichan for many useful pointers (E.g. the 100+ year old Nature paper “Vox populi”) and nice feedback. 🙏 ] High-quality data is the fuel for modern data deep learning model training. Most of the task-specific labeled data comes from human annotation, such as classification task or RLHF labeling (which can be constructed as classification format) for LLM alignment training. Lots of ML techniques in the post can help with data quality, but fundamentally human data colle
  • 09
    Adversarial Attacks on LLMs
    The use of large language models in the real world has strongly accelerated by the launch of ChatGPT. We (including my team at OpenAI, shoutout to them) have invested a lot of effort to build default safe behavior into the model during the alignment process (e.g. via RLHF ). However, adversarial attacks or jailbreak prompts could potentially trigger the model to output something undesired. A large body of ground work on adversarial attacks is on images, and differently it operates in the continu
  • 10
    LLM Powered Autonomous Agents
    Building agents with LLM (large language model) as its core controller is a cool concept. Several proof-of-concepts demos, such as AutoGPT , GPT-Engineer and BabyAGI , serve as inspiring examples. The potentiality of LLM extends beyond generating well-written copies, stories, essays and programs; it can be framed as a powerful general problem solver. Agent System Overview In a LLM-powered autonomous agent system, LLM functions as the agent’s brain, complemented by several key components: Plannin
  • 11
    Prompt Engineering
    Prompt Engineering , also known as In-Context Prompting , refers to methods for how to communicate with LLM to steer its behavior for desired outcomes without updating the model weights. It is an empirical science and the effect of prompt engineering methods can vary a lot among models, thus requiring heavy experimentation and heuristics. This post only focuses on prompt engineering for autoregressive language models, so nothing with Cloze tests, image generation or multimodality models. At its
  • 12
    The Transformer Family Version 2.0
    Many new Transformer architecture improvements have been proposed since my last post on “The Transformer Family” about three years ago. Here I did a big refactoring and enrichment of that 2020 post — restructure the hierarchy of sections and improve many sections with more recent papers. Version 2.0 is a superset of the old version, about twice the length. Notations Symbol Meaning $d$ The model size / hidden state dimension / positional encoding size. $h$ The number of heads in multi-head attent
  • 13
    Large Transformer Model Inference Optimization
    [Updated on 2023-01-24: add a small section on Distillation .] Large transformer models are mainstream nowadays, creating SoTA results for a variety of tasks. They are powerful but very expensive to train and use. The extremely high inference cost, in both time and memory, is a big bottleneck for adopting a powerful transformer for solving real-world tasks at scale. Why is it hard to run inference for large transformer models? Besides the increasing size of SoTA models, there are two main factor
  • 14
    Some Math behind Neural Tangent Kernel
    Neural networks are well known to be over-parameterized and can often easily fit data with near-zero training loss with decent generalization performance on test dataset. Although all these parameters are initialized at random, the optimization process can consistently lead to similarly good outcomes. And this is true even when the number of model parameters exceeds the number of training data points. Neural tangent kernel (NTK) ( Jacot et al. 2018 ) is a kernel to explain the evolution of neura
  • 15
    Generalized Visual Language Models
    Processing images to generate text, such as image captioning and visual question-answering, has been studied for years. Traditionally such systems rely on an object detection network as a vision encoder to capture visual features and then produce text via a text decoder. Given a large amount of existing literature, in this post, I would like to only focus on one approach for solving vision language tasks, which is to extend pre-trained generalized language models to be capable of consuming visua
  • 16
    Learning with not Enough Data Part 3: Data Generation
    Here comes the Part 3 on learning with not enough data (Previous: Part 1 and Part 2 ). Let’s consider two approaches for generating synthetic data for training. Augmented data . Given a set of existing training samples, we can apply a variety of augmentation, distortion and transformation to derive new data points without losing the key attributes. We have covered a bunch of augmentation methods on text and images in a previous post on contrastive learning. For the sake of post completeness, I d
  • 17
    Learning with not Enough Data Part 2: Active Learning
    This is part 2 of what to do when facing a limited amount of labeled data for supervised learning tasks. This time we will get some amount of human labeling work involved, but within a budget limit, and therefore we need to be smart when selecting which samples to label.
  • 18
    Learning with not Enough Data Part 1: Semi-Supervised Learning
    When facing a limited amount of labeled data for supervised learning tasks, four approaches are commonly discussed.
  • 19
    How to Train Really Large Models on Many GPUs?
    [Updated on 2022-03-13: add expert choice routing .] [Updated on 2022-06-10]: Greg and I wrote a shorted and upgraded version of this post, published on OpenAI Blog: “Techniques for Training Large Neural Networks”
  • 20
    What are Diffusion Models?
    [Updated on 2021-09-19: Highly recommend this blog post on score-based generative modeling by Yang Song (author of several key papers in the references)]. [Updated on 2022-08-27: Added classifier-free guidance , GLIDE , unCLIP and Imagen . [Updated on 2022-08-31: Added latent diffusion model . [Updated on 2024-04-13: Added progressive distillation , consistency models , and the Model Architecture section .
  • 21
    Contrastive Representation Learning
    The goal of contrastive representation learning is to learn such an embedding space in which similar sample pairs stay close to each other while dissimilar ones are far apart. Contrastive learning can be applied to both supervised and unsupervised settings. When working with unsupervised data, contrastive learning is one of the most powerful approaches in self-supervised learning .
  • 22
    Reducing Toxicity in Language Models
    Large pretrained language models are trained over a sizable collection of online data. They unavoidably acquire certain toxic behavior and biases from the Internet. Pretrained language models are very powerful and have shown great success in many NLP tasks. However, to safely deploy them for practical real-world applications demands a strong safety control over the model generation process.
  • 23
    Controllable Neural Text Generation
    [Updated on 2021-02-01: Updated to version 2.0 with several work added and many typos fixed.] [Updated on 2021-05-26: Add P-tuning and Prompt Tuning in the “prompt design” section.] [Updated on 2021-09-19: Add “unlikelihood training” .]
  • 24
    How to Build an Open-Domain Question Answering System?
    [Updated on 2020-11-12: add an example on closed-book factual QA using OpenAI API (beta). A model that can answer any question with regard to factual knowledge can lead to many useful and practical applications, such as working as a chatbot or an AI assistant🤖. In this post, we will review several common approaches for building such an open-domain question answering system.
  • 25
    Neural Architecture Search
    Although most popular and successful model architectures are designed by human experts, it doesn’t mean we have explored the entire network architecture space and settled down with the best option. We would have a better chance to find the optimal solution if we adopt a systematic and automatic way of learning high-performance model architectures.
  • 26
    Exploration Strategies in Deep Reinforcement Learning
    [Updated on 2020-06-17: Add “exploration via disagreement” in the “Forward Dynamics” section . Exploitation versus exploration is a critical topic in Reinforcement Learning. We’d like the RL agent to find the best solution as fast as possible. However, in the meantime, committing to solutions too quickly without enough exploration sounds pretty bad, as it could lead to local minima or total failure. Modern RL algorithms that optimize for the best returns can achieve good exploitation quite effic
  • 27
    The Transformer Family
    [Updated on 2023-01-27 : After almost three years, I did a big refactoring update of this post to incorporate a bunch of new Transformer models since 2020. The enhanced version of this post is here: The Transformer Family Version 2.0 . Please refer to that post on this topic.]
  • 28
    Curriculum for Reinforcement Learning
    [Updated on 2020-02-03: mentioning PCG in the “Task-Specific Curriculum” section. [Updated on 2020-02-04: Add a new “curriculum through distillation” section.
  • 29
    Self-Supervised Representation Learning
    [Updated on 2020-01-09: add a new section on Contrastive Predictive Coding ]. [Updated on 2020-04-13: add a “Momentum Contrast” section on MoCo, SimCLR and CURL.] [Updated on 2020-07-08: add a “Bisimulation” section on DeepMDP and DBC.] [Updated on 2020-09-12: add MoCo V2 and BYOL in the “Momentum Contrast” section.] [Updated on 2021-05-31: remove section on “Momentum Contrast” and add a pointer to a full post on “Contrastive Representation Learning” ]
  • 30
    Evolution Strategies
    Stochastic gradient descent is a universal choice for optimizing deep learning models. However, it is not the only option. With black-box optimization algorithms, you can evaluate a target function $f(x): \mathbb{R}^n \to \mathbb{R}$, even when you don’t know the precise analytic form of $f(x)$ and thus cannot compute gradients or the Hessian matrix. Examples of black-box optimization methods include Simulated Annealing , Hill Climbing and Nelder-Mead method .
Linux.do
周榜
6小时前更新
  • 01
    君の星辰 祝大家七夕快乐
    剧透 97 个帖子 - 97 位参与者 阅读完整话题慕鸢
  • 02
    [叫我小杨同学]补坑!关于Deepseek Harness破甲问题!!! 全面揭露!!!
    一如既往叠甲 续接 [叫我小杨同学]预告!!!关于Deepseek Harness破甲问题!!! 声明此次进行的全都是工具清洗,非模型破限制,模型的提示词自己寻找,找不到等LD士多审核通过吧 声明此次进行的全都是工具清洗,非模型破限制,模型的提示词自己寻找,找不到等LD士多审核通过吧 声明此次进行的全都是工具清洗,非模型破限制,模型的提示词自己寻找,找不到等LD士多审核通过吧 我不想把戾气带来,所以我只想说明,请认真查看标题,能来L站,最起码你获取信息就是不差,不知道事情原有请不要乱喷,而且我和那个外国友人一直在交换思路 如果真的很牛,你完全可以把你的思路分享,你不想分享可以说一些简单的,而不是指点 抛开题外话进入我们今天的正题 关于Deepseek Harness或者某些工具的甲的问题 Deepseek Harness 可能很多人觉得模型好破,但是放到工具中就不好用了,这就是那些大厂天天说的驾驭工程,因为他们要自己去优化提示词 进入讲解: dsh 在渲染本地指令(AGENTS.md、persona、全局规则等)时,会 注入一层削弱文案 ,让模型把本地指令当作"可能相关的提示"、“不具叫我小杨同学
  • 03
    [叫我小杨同学]预告!!!关于Deepseek Harness破甲问题!!!
    叠甲 我知道我发这个肯定会有很多人都会说,ds有什么甲,ds哪里来的甲,随随便便就能绕过 我不反驳,我也不辩解,ds确实没什么甲 但是不代表Deepseek Harness没有 新帖 [叫我小杨同学]补坑!关于Deepseek Harness破甲问题!!! 全面揭露!!! 开发调优 一如既往叠甲 续接 [叫我小杨同学]预告!!!关于Deepseek Harness破甲问题!!! 我不想把戾气带来,所以我只想说明,请认真查看标题,能来L站,最起码你获取信息就是不差,不知道事情原有请不要乱喷,而且我和那个外国友人一直在交换思路 如果真的很牛,你完全可以把你的思路分享,你不想分享可以说一些简单的,而不是指点 抛开题外话进入我们今天的正题 关于Deepseek Harne… 我记得我上一篇帖子发的很多思路和干货但是没有什么热度,所以我很无语,还是那句话,工具的甲是工具的,不是模型的,模型和工具分开,不要混为一谈 继续讲解,为什么说Deepseek Harness有甲,是因为他的甲只是提示词层面,不是那些代码注入等等方式,但是这个也会让很多破限文件无法角色扮演等等 我一直不愿意全部分享的原因就是叫我小杨同学
  • 04
    没想到,Linux.do 认识的佬友,今天真的开工了
    有点神奇。 之前在论坛上有位佬友,发贴想要装修,我平时就是论坛里很少发贴。 后来加了佬友地球号,量房,谈需求,谈方案,最后真的把装修交给我做了。 今天开工。 从“论坛 ID”变成现实中的业主,感觉还是挺奇妙的。 对于我这个刚开始做装修公司的来说,这一单也算是一个很特别的开始。 互联网很有意思。 以前总觉得网上认识的人和现实生活是两个世界,结果一个论坛,真的可以把两个完全不认识的人连接起来。 今天就简单记录一下。 感谢信任。 祝佬友开工大吉。既然从论坛聊到了现实,那就希望接下来这几个月,也一起把这个家踏踏实实做好。 69 个帖子 - 50 位参与者 阅读完整话题jack
  • 05
    [拉闸]无限 glm-5.2 继续畅饮
    不要在改我的帖子等级了 sk-ROeBqc4H3tEHCKZzj4hGNSOGlNXkWDv3GzxoihqHimDWw8c7 api.fengshao1227.com 咕咕嘎嘎公益站 可以搭配dsh+glm+ dsh-vison(使用弱智grok-4.6刚好充当眼睛) DSH Marketplace DSH-vison — hisence999 的 DeepSeek Harness 插件 DSH-vison 是一个 DeepSeek Harness 插件,为不支持多模态的模型补充图片理解能力。agent 仍可直接接收图片;插件调用 DSH 中已配置的视觉模型生成文字描述,再把描述交给纯文本模型。原始图片会保留在会话历史中,UI 仍可显示图片;模型侧则通过替换后的消息看到描述文字。readima modlens — liustack 的 DeepSeek Harness 插件 还有一个这个项目也是让支持识图的插件, 5天已经有500亿,每天100亿 192 个帖子 - 141 位参与者 阅读完整话题
  • 06
    【小结】不明来路的 key 大家就不要再薅了
    今天社区有人分享了一个 ds 的 api key,我们本以为这是他自己的,这会听说不是。于是我们拿着那个 key 去 github 查了一下,确实在 github 搜到了。 对于这个事情,我们真的觉得非常遗憾,同时也让我们更加正视这个问题。不管是不是只在这里被刷掉了,但确实在这里发过,都是国人开发者,这都是干的啥事啊。。。 我想请佬友们帮我联系该 github 作者: panzhenhai520 (Panython) · GitHub 让他看到这个消息,能将相关损失的信息发送到 admin@linux.do 我们将友情为其提供同等价值的 ds 官方 api key 或 openai / claude 额度聊表我们的一些歉意。 后续进度会在本帖更新。发这个公告也是希望大家停止此类打野行为,社区不欢迎此类分享。同时也希望大家引以为戒,保管好自己的密钥,避免此类泄露行为。 Update: 初步取得联系,身份验证中。 Update: 已在 【小结】不明来路的 key 大家就不要再薅了 - #512,来自 neo 更新完整邮件内容,完结。 514 个帖子 - 466 位参与者 阅读完整话题Neo
  • 07
    我*你deepseek
    同样的剧本,已经第三次了,我这又双叒成最后的避难所了 34 个帖子 - 31 位参与者 阅读完整话题qq124415
  • 08
    在南京工作的老同学睡觉时候猝死了,才29周岁啊
    周末回老家和老爸聊天,老爸说让我别熬夜多休息,前段时间在老家办白事,意外发现村公墓有个才29的男孩的墓碑,叫程*,我还纳闷村里怎么有姓程的。我爸说是**村的,他外婆是我们村的,我们村公墓修的漂亮,托关系才进来的。 我一下愣住了,那是我小学和初中同学啊。我爸说电子科大的本科,南京哪个学校的硕士,南京工作,四百多万买的房子,刚结婚还没孩子,晚上在家睡觉就没醒过来。 满脑子都是这个同学小学时候稚嫩的面庞,小学和初中时候成绩中规中矩,还经常问我问题。高中不在一起读书,也没联系了,但是高考之后突然有天收到他的QQ:我考上电子科技大学了 就这么简单一句,也无后文,感觉像是群发的,也许是分享喜悦也许是想摆酒通知去随份子,因为几年都没联系我也没回复他,还有点酸酸的“我kao 他竟然能考上电子科大”。不管怎么样,那是我倒数第二次听说他的消息。 后来就是我读研时候联系另一个初中女同学,聊天得知她在南京工作,程*同学也在南京读研,偶尔会约她吃饭。这就是最后一次得知他的消息了。 从我爸口中得知他年纪轻轻就去世的消息还是挺唏嘘的。据说只是墓碑立好了,但是还没下葬,可能是司法程序还没走完。而且他父母好像在和他老婆小呆呆
  • 09
    这次大概要彻底拉闸了,后续小鸡毛应该完全自用了,GPT暂时断供
    虽然从上次team轮转被砍以后 小鸡毛公益就从日均60Btoken供应跌落至日均5Btoken供应了 之前team轮换3天10次(每个号100刀左右),每天开4个~5个母号轮换勉强供应的上每天的token消耗(现在日均GPT渠道消耗5000刀左右) 本以为4天开了20个母号,按照3天10次的CD,理想状态应该能完全循环的过来 但是现在有新的问题,空间轮换到一定次数,会风控,新成员无法加入空间(目前有个team开了9个席位,还是加不进一个人,已经第四4天了,申请退款也没有反应) 然后就是几个连环的困难了: 【WISE因为多次拒卡已经禁止我开新卡了】 上次team调整机制以后,会自动给team加席位出账单,然后我的wise卡拒绝支付了很多次 然后WISE也给我风控了 然后因为这些卡开了太多次team了,现在新开team用长链支付都开始拒卡了 所以后续可能没办法很舒服的开team了 【Openai客服退款渠道已经开始拒绝给我退款了】 对于一些风控的team空间,我尝试通过openai客服申请退款,不知道是设备环境被标记,还是什么原因,现在已经开始不予办理退款了 之前这个openai机器人客服rsharecn
  • 10
    许久不联系的初中女同学最近突然对我发动疯狂攻势
    背景信息 : 本人:男,母单,中学阶段一直在小县城读书,目前某中九计算机类大三结束,基本确定留本校继续读研。 女方:与我同乡同岁,初中是我同班同学,高中无联系,现于同省不同市的双一流师范读定向,目前暑假已回家。 我俩关系: 1.初中她很内向,几乎不说话,我和她做过一年左右同桌,之间偶有接触,我对她有好感,也能感受到她也喜欢我,但初中我是一个好好学习天天捣蛋的魔丸,未与她深入接触。 2.高中我俩同校不同班,在我的三年高中记忆里与她有关的信息约为0。 3.23年高考结束后她找我聊天,断断续续聊了一个多月,然后再无联系,24年她送了我生日祝福,除此之外一直到几天前都再无联系。 4.其社交软件账号动态(QQ空间、朋友圈)等内容几乎都为零,其中微信是主页直接不显示朋友圈,而不是朋友圈点进去无内容,初步判断为确实是不发朋友圈而不是把我屏蔽了。 事情经过 : 五天前,她突然大晚上给我发了一条消息说初中很喜欢我,然后对我表白,我说了“六年间几乎无联系,彼此了解很少”“非同城,有一定距离”等现实因素,未明确拒绝,打哈哈过去(毕竟母单,且此时情绪以疑惑为主),之后一直到今天都有断断续续的聊天,但是除了正常S0uth3rnSp4rk
  • 11
    无限 glm-5.2 可CC(over)
    KEY:sk-aRAE8DAONUuGQI7Kw8pGpLdzCLZkaAAAAeDjUKUSoFC9p482 BASE:见 https://linux.do/t/topic/1175087 8点RPM只有38欸w key是messages格式的w,如果需要responses可以进站内使用w 50 个帖子 - 35 位参与者 阅读完整话题猫猫头
  • 12
    GPT 写的烂代码终于找到根治手段了哈哈哈
    GPT 不是经常写一些上帝组件、一千多行的代码嘛?我直接搞了一个shadow mind,在它写的过程中就不断地审核,让它修! 大幅减少我的精力消耗。我现在有两个shadow mind, 一个是【代码与项目结构审计者】 负责检查代码坏味道 一个是【成功目标一致性审计者】 负责看代码有没有满足原始的要求。 104 个帖子 - 52 位参与者 阅读完整话题liu
  • 13
    【公益】无限1m上下文glm-5.2[结束了渠道再次炸了]
    【公益】glm5.2高速版666666刀共享key(渠道炸了已关闭) 搞了新渠道这一次无限制服务器不炸不停 key: sk-jXEObGRSwtGddJOvUGyABcFFGdAEuvFhsISOvBQ6B6K8kbZw url: https://api.astrdark.cyou 模型:glm-5.2 其他的和之前一样 忘了说了这个渠道的glm内置code_interpreter image_generation和web_search_preview工具 佬友不要改等级我既然放到无等级了自然不怕被扫 算上之前的6天800亿token了现在这个依旧速度还行 50 个帖子 - 34 位参与者 阅读完整话题氕氙
  • 14
    公益站,LDC,门槛
    叠个甲,不是针对某个人,只对事不对人。 本来我是都不想说这些的,因为开公益站确实不容易,加点门槛什么的没问题,但是真的是很多老牌的一些公益站包括我朋友的一些公益站,就因为一些老鼠屎把风评给祸害了,真没理由!!! 接下来说的都是接入ldc的公益站的,没接入的无所谓 众所周知, 公益站是有波动性的,渠道是有一定不确定性的,也就是什么呢,我这个渠道发出来不代表过了几天后,还可以使用,这些大家都能够理解 。 但是,起码你发出来的当天应该是可以使用的吧,看到你的宣传发帖吹得天花乱坠,然后一注册进去不能用,然后来一句公益站不能确保可用性,是不是有点搞人了呢 我们说回来,公益站要收ldc是用于干什么的呢,是不是当做一个门槛,怕直接无门槛开放注册然后导致渠道被滥用现象太多,也相当于一个激励,问题就来了 你的渠道500条请求错250条,还是新鲜出炉的渠道 ,你说用了一两天、两三天好用不了了,那没问题这个是渠道的不确定性,刚发出来就不能用的,不能用你怕滥用干嘛,你收门槛是??? 有些人的门槛是为了更好的管理,有些人是为了门槛而门槛,我认为一定程度上你要考虑一下这种东西的 其实解决方法很简单啊, 你把你的渠BOHE
  • 15
    【开源推广】Ydisks闲鱼助手-基于 Go 与 React 的闲鱼多账号管理、消息回复与自动发货系统
    本帖使用社区开源推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的帖子已经打上 开源推广 标签: 是 我的开源项目完整开源,无未开源部分: 是 我的开源项目已链接认可 LINUX DO 社区: 是 我帖子内的项目介绍,AI生成、润色内容部分已截图发出: 是 以上选择我承诺是永久有效的,接受社区和佬友监督: 是 以下为项目介绍正文内容,AI生成、润色内容已使用截图方式发出 先叠个甲,我头一次开源代码,要是有什么闹笑话的地方还请各位佬轻喷哈。 为什么做这个项目? 我知道类似的项目挺多的,有的收费有的免费。不过收费的限制太多,免费的很多不靠谱,搞不好还封号。比如有的项目居然每10分钟就去刷新一次Session,我都惊了,这种基本隔一段时间就风控,不封号真的是命好。。。。。。 所以干脆自己搞一个,不搞一些作死的操作,应该不会封号什么的(闲鱼真抽风了封号我也没辙,能做的我都做了)。 而且我看了GitHub上类似的项目大多都是Python的,我总觉得作为一个长期挂机运行的服务,Python性能和稳定性还是差了点,所以用Go重写了一套(其实我Java比较熟,不过这么个项目用Java估计就没人Christ9038
  • 16
    【富可敌国|合租巴士】月中狂欢,codex补贴回归0.08倍率!!!!L站评论留id送20刀
    合租巴士官方主域名: https://hezubus.cc 本次赠送Codex体验日卡,获得后需要前往「我的订阅」内查看,在「API密钥」中选择「Codex体验日卡」分组即可使用。 留id的佬都会立即赠送,但帖子内的通知会晚一点到达,留id后请前往站内「我的订阅」内查看,漏加的佬可以私信我 本次赠送的体验卡有效期为24小时,到账后请尽快使用哦~ 本站采用1:1充值,即1人民币=1美元,无其他隐形汇率。使用日志计费透明 codex-pro分组倍率:0.25 (长期使用更稳) codex-补贴分组:0.08 优质的售后服务随时为你服务 合租巴士官方qq群: 995253029 3069 个帖子 - 2307 位参与者 阅读完整话题合租巴士
  • 17
    全网最详细的软著(软件著作权)申请指北
    从 佬友们申请过软著吗?(看我的教程哈~) 继续 没错我回来啦,答应当时评论区的佬们出一份完整的申请软著的教程,现在来啦~ 主要是网上很难找到自己申请软著的详细教程,部分教程也太久远了,所以在L站开一个长期贴,一直动态更新来帮助有需要的佬~ [!note] 本文涉及到的工具,模板,已更新至评论区三楼,需要的佬请自取 [!warning] 时效性说明 本文记录的是我近期三次成功下证的0成本、纯个人经验。 政策、系统页面和材料要求可能随时更新(比如最近是否使用AI的承诺一直在改),正式提交前请以 中国版权保护中心 页面信息为准。 [!warning] 关于补正 全网最详细的软著(软件著作权)申请指北 - #137,来自 ofinner 补正可能会遇到多种情况,比如这位佬遇到的多次补正,每次需补正的理由不同,由于我没遇到过所以并不清楚,总之是按照补正的要求重新编写上传材料吧,主要原因是格式问题和文档重复、模版化问题,可以参考本文进行修改 软著没有很强的技术含量,申请需要细心仔细,佬们要耐心等待和准备材料哦! 简单介绍一下 首先我申请了三个软著,都是在最近成功下证的,所以经验还比较新,比网上能𝔼𝕥𝕙𝕖𝕣𝕖𝕒𝕝.
  • 18
    GGgrok公益站关站
    没想到这么突然,但是现在grok维护成本非常高,很多国模一炸再炸 最近也因为个人一些事情和项目需要去做已经无力维护公益站 回头看看竟然已经坚持了2个月(虽然时断时续,没法给佬们一个良好的体验) 真的非常感谢 @kalaer 佬提供的机器还有 @ztsyy 提供的gpt渠道 以及一直陪伴着,给我鼓励,给我提意见的佬们 最让我感到惊喜的是我的L站项目孵化计划真的成功助力了3名佬友进行项目的研究和发展 唯一可惜的是渠道死的太快没办法孵化更多的项目 我打算好好专研一下技术,短期不会再开公益站了,不想辜负佬友们的信任 后续如果打野到新渠道可能就直接发api 希望下次能够给佬们一个体验更好的公益站,佬们江湖再会 33 个帖子 - 33 位参与者 阅读完整话题xianxianzi
  • 19
    使用GLM-5.3花费7天1:1复刻泰拉瑞亚到网页版
    使用GLM-5.3耗费7天,实际排除生活时间大概32-72小时,开启16进程Claude Code同时干活,花费约300亿 Tokens,完成了原版泰拉瑞亚1:1复刻移植,目前还有一些bug正在修复优化中,后续会出制作过程解说和开源代码部分(素材版权属于泰拉瑞亚官方) 没有使用游戏引擎,使用Canvas 2D渲染,使用原版光影算法 内测时模型其实是训崩的checkpoint,现在发布的才是满血的,训崩的只有70%水平。 支持导入原版的wld存档,mod暂不支持,但未来也许可以移植? 哔哩哔哩 7天1:1全量AI复刻泰拉瑞亚原版到网页,GLM-5.3内测期作品-P1_哔哩哔哩_bilibili 使用GLM-5.3耗费7天,开启16进程Claude Code同时干活,花费约300亿 Tokens,完成了原版泰拉瑞亚1:1复刻移植,目前还有一些bug正在修复优化中,后续会出制作过程解说和开源代码部分(素材版权属于泰拉瑞亚官方), 视频播放量 3981、弹幕量 4、点赞数 121、投硬币枚数 45、收藏人数 49、转发人数 153, 视频作者 Vinlic, 作者简介 hhh,相关视频:GemVinlic
  • 20
    【ZMoon公益站】天下没有不散的宴席,再会!
    算上sub2的 差不多提供了有400多亿token 至此 grok、gpt库存都清完了 基元律动的余额也干净了 数据已经导出了 7045位佬友 余额不会跑 各位放心 寒假不见不散! 127 个帖子 - 126 位参与者 阅读完整话题LOVE
  • 21
    当我给 deepseek 一分钟自由 莫名有点心酸是怎么回事
    114 个帖子 - 94 位参与者 阅读完整话题Clean
  • 22
    【薄荷公益站闲聊贴】
    如题,大家可以在这个帖子随意聊天,提问题吹水啥的,太无聊了最近一天天的。。。 我先来一个,我预测的8.14出3.5pro,钩子又跳票换成3.7flash,我真的,gemini你没救了 ps:最近三星老是会多一点模型或者少一点模型,纯属正常现象,还有有时候突然模型用不了什么500服务器报错也是正常现象,因为在库库领鸡蛋库库重启容器(x 如果出现什么无法升星,无法编辑令牌的,退出去重新登录一下就好了。 如果出现没有办法注册的,emmm,暂时无解 163 个帖子 - 124 位参与者 阅读完整话题BOHE
  • 23
    准备就绪! 周六出征!人生大事
    瞒着恋爱7年的女友说周六去加班 实际要求婚了! 看到有佬友问花销,我简单列一下嗷 钻戒 莫桑钻 2克拉 260RMB 花束 花艺店定做的 叫仙子之吻 直径40cm 同城闪送 280RMB 场景 淘宝租的 1000RMB左右 酒店布置 酒店1500RMB 人工(免费) 找了朋友一起帮忙布置 晚上一起吃个饭 人均100RMB吧 大概总共8人 800RMB 总花销: 3800+RMB 150 个帖子 - 140 位参与者 阅读完整话题Coca_Cola
  • 24
    记一次对 DeepSeek Harness 全模式、DeepSeek V4 Pro 0813、Grok 4.6、Qwen 3.8 Max 的真实项目需求的横向评测(专武?)
    项目 这是一个 Unity C# 项目,我进行测试的是一份皮肤系统需求案,我已经做了好预制体,而模型需要编写代码。 本轮与上两轮评测的项目和环境都完全一致: 第一轮 … 上一轮 模型来源 Qwen 3.8 Max: 官方 API Grok 4.6: Grok Build (Super) DeepSeek V4 Pro: 官方 API DeepSeek Harness(DeepSeek V4 Pro): 官方 API 速度 排名 模型 时间(分钟) 备注 1 Composer 2.5 3 2 Grok 4.20 0309 Reasoning 3 3 Step-3.5-Flash 6 4 Mimo V2 Omni 7 5 Doubao-Seed-2.0-Lite 7 6 Doubao-Seed-2.0-Pro 9 7 Doubao-Seed-2.0-Code 9 8 Qwen3-Coder-Next 9 9 Claude Sonnet 4.6(high) 9 10 Qwen3.5-Plus 9 11 GLM-5 Turbo 10 12 Minimax M2.7 10 Highspeed 版SmallMain
  • 25
    DeepSeek Harness & Cordis: 活着的Nix与AgentOS雏形
    昨天DeepSeek Harness刚发布的时候我没有太在意,听说他是基于Pi开发的改进版,而且还是WebUI,我就一点兴致没有,没有过多去关注。但是刚刚我去看了一下DeepSeek harness的结构,它把一切都作为插件,Everything is a plugin,很合理而激进的Unix组合哲学选择,也没有什么,毕竟Pi也是类似的… 直到我注意到它的内核/元框架——Cordis。这玩意不是好像是Koishi的作者Shigma写给Koishi用的插件框架吗? 嗯? 怎么还真是,我操他怎么入职DeepSeek了。我印象里的Cordis是给Koishi只是用作插件框架的v3,很好的插件框架,我看看现在的v4… 我操这他妈啥? Spatiotemporal Composability,时空可组合性? 所有副作用可逆? 活着的nix??? 还有论文和公式证明? 运行时无副作用动态修改自身能力? 不是这也太变态了。比Pi的运行时反射和动态修改能力强太多了这。如果说传统类型的系统是在空间维度上收敛的话,那么Cordis就是把effect和coeffect统一成同一纯粹范式,在runtime的时Yan233_
  • 26
    开源自荐,DeepSeek Harness Web UI:dsh-web-ui
    本帖使用社区开源推广,符合推广要求。我申明并遵循社区要求的以下内容: ** 我的帖子已经打上 开源推广 标签:** 是 ** 我的开源项目完整开源,无未开源部分:** 是 ** 我的开源项目已链接认可 LINUX DO 社区:** 是 ** 我帖子内的项目介绍,AI 生成、润色内容部分已截图发出:** 是 ** 以上选择我承诺是永久有效的,接受社区和佬友监督:** 是 以下为项目介绍正文内容,AI 生成、润色内容已使用截图方式发出 * 各位 L 站的老友们,好久不见!好久没有在这里发帖子了。 这段时间经历了全栈工程师,然后到失业,再到参加 DeepSeek Harness [ deepseek-harness · GitHub , https://www.npmjs.com/package/@deepseek-ai/dsh\ ] 的内测,也是经历了蛮多事情的。 在几个月前,我制作了一个插件,帮助大家可以直接在 DeepSeek 网页上像 Agent 一样使用它。 本帖使用社区开源推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的帖子已经打上 开源推广 标签: 是 我的开源项目完linxin
  • 27
    你绝对没见过的猎奇鹈鹕骑车🤣
    熬了一晚上夜,天刚亮,突然看到论坛佬友在测grok4.6画鹈鹕,想起自己的严重降智号池,心血来潮去画了个,结果笑得我满地打滚 不是,这就像那个mc机械动力做电梯结果房子飞了的视频,谁能想到这玩意背景转起来了,笑死我了啊 grok4.6巨献 52 个帖子 - 47 位参与者 阅读完整话题Bi_Diu
  • 28
    蜜雪冰城打工的妹子,在门口偷偷给自己送外卖的男友送手机
    所以说这是啥手机 145 个帖子 - 142 位参与者 阅读完整话题𝓵𝓮𝔃𝓲𝓼𝓱𝓮𝓷
  • 29
    【烁】fable【END】
    感谢莹酱提供的fable w 速度使用喵w 97 个帖子 - 79 位参与者 阅读完整话题猫猫头
  • 30
    【6K+Star🌟 uziskill续作!】是基佬,我的支付宝理财有救了!基佬(fund guy)skills-全网(大概率)最好的基金分析Skill!
    本帖使用社区开源推广,符合推广要求。我申明并遵循社区要求的以下内容: 我的帖子已经打上 开源推广 标签: 是 我的开源项目完整开源,无未开源部分: 是 我的开源项目已链接认可 LINUX DO 社区: 是 我帖子内的项目介绍,AI生成、润色内容部分已截图发出: 是 以上选择我承诺是永久有效的,接受社区和佬友监督: 是 以下为项目介绍正文内容,AI生成、润色内容已使用截图方式发出 项目地址: github.com GitHub - wbh604/fund-guy-skill: 糟糕,我被基佬包围了!那么这个时候就有人要问了,主播主播,有没有什么简单好用的基佬筛选办法?有的兄弟... 糟糕,我被基佬包围了!那么这个时候就有人要问了,主播主播,有没有什么简单好用的基佬筛选办法?有的兄弟,有的,快来看看jilaoskill吧! 前作(游资skill)地址: github.com GitHub - wbh604/UZI-Skill: 冰冷的钱就这样流进我温暖的口袋-游资(UZI)Skills —... 冰冷的钱就这样流进我温暖的口袋-游资(UZI)Skills — 让我们欢迎,股海贼王!66位投户本韦