全部/科技/Stanford CRFM · Blog

技术博客与 newsletter feed · Stanford CRFM · Blog

HISTORY近 30 天历史柱高表示当天去重热搜数量
1007
09/08—10/07 有历史数据
  • 01
    HELM Arabic Enterprise
    We present HELM Arabic Enterprise, a leaderboard for transparent, reproducible evaluation of large language models on Arabic-language benchmarks designed around enterprise use cases. The leaderboard was developed in collaboration with Arabic.AI and builds on the HELM evaluation methodology: standardized prompting, fully logged requests and responses, and reproducible scoring through the open-source HELM framework.Yifan Mai
  • 02
    HELM Arabic
    As part of our efforts to better understand the multilingual capabilities of large language models (LLMs), we present HELM Arabic, a leaderboard for transparent and reproducible evaluation of LLMs on Arabic language benchmarks. This leaderboard was produced in collaboration with Arabic.AI.Yifan Mai
  • 03
    HELM Long Context
    We introduce the HELM Long Context leaderboard for transparent, comparable and reproducible evaluations of long context capabilities of recent models. Introduction Recent Large Language Models (LLMs) support processing long inputs with hundreds of thousands or millions of tokens. Long context capabilities are important for many real-world applications, such as processing long text documents, conducting long conversations or following complex instructions. However, support for long inputs does noYifan Mai
  • 04
    Reliable and Efficient Amortized Model-Based Evaluation
    TLDR: We enhance the reliability and efficiency of language model evaluation by introducing IRT-based adaptive testing, which has been integrated into the HELM framework.Sang Truong
  • 05
  • 06
    BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
    We introduce BountyBench , a benchmark featuring 25 systems with complex, real-world codebases, and 40 bug bounties that cover 9 of the OWASP Top 10 Risks. Key Takeaways BountyBench is a benchmark containing 25 diverse systems and 40 bug bounties, with monetary awards ranging from $10 to $30,485, covering 9 of the OWASP Top 10 Risks . It is designed to evaluate offensive and defensive cyber-capabilities in evolving real-world systems. To capture the vulnerability lifecycle from discovery to repaAndy K. Zhang
  • 07
    HELM Capabilities: Evaluating LMs Capability by Capability
    Introducing HELM Capabilities, a benchmark that evaluates language models across a curated set of key capabilities, providing a comparison of their strengths and weaknesses. Evaluating language models is a dynamic and critical process as models continue to improve rapidly. Understanding their strengths and weaknesses is essential for external users to determine which models suits their needs. Two years ago, we introduced the Holistic Evaluation of Language Models (HELM) as a framework to assessJialiang Xu
  • 08
    General-Purpose AI Needs Coordinated Flaw Reporting
    Today, we are calling for AI developers to invest in the needs of third-party, independent researchers, who investigate flaws in AI systems. Our new paper advocates for a new standard of researcher protections, reporting and coordination infrastructure. The paper, In House Evaluation Is Not Enough: Towards Robust Third-Party Flaw Disclosure for General-Purpose AI, has 34 authors with expertise in machine learning, law, security, social science, and policy. Introduction Today, we are calling forShayne Longpre
  • 09
  • 10
    Advancing Customizable Benchmarking in HELM via Unitxt Integration
    The Holistic Evaluation of Language Models (HELM) framework is an open source framework for reproducible and transparent benchmarking of language models that is widely adopted by academia and industry. To meet HELM users’ needs for more powerful benchmarking features, we are proud to announce our collaboration with Unitxt, an open-source community platform developed by IBM Research for data preprocessing and benchmark customization. The integration of Unitxt into HELM gives HELM users access toYifan Mai