awesome-ai-leaderboard

作者 SAILResearch已验证

A curated list of awesome leaderboard-oriented resources for AI domain

379
Stars
58
Forks
2026/8/23
添加时间

⚠️ 第三方软件声明

本 Skill 为第三方开源软件,独立托管于 GitHub。SkillTip 仅为信息目录,不控制或维护底层仓库。所显示的安全检查为自动化且范围有限,安装前请自行审查源码。

阅读服务条款

安装

添加到你的 Claude Code skills 目录:

# Add to your Claude Code skills
git clone https://github.com/SAILResearch/awesome-ai-leaderboard

快速入门

使用 awesome-ai-leaderboard 等 Skills 的指南。

安全报告

已验证

上次扫描:—

{
  "status": "PASSED",
  "issues": []
}

README.md

lbops

Awesome AI Leaderboard

Awesome AI Leaderboard is a curated list of awesome AI leaderboards, along with various development tools and evaluation organizations according to our recent survey:

On the Workflows and Smells of Leaderboard Operations (LBOps):
An Exploratory Study of Foundation Model Leaderboards

Zhimin (Jimmy) Zhao, Abdul Ali Bangash, Filipe Roseiro Côgo, Bram Adams, Ahmed E. Hassan

Software Analysis and Intelligence Lab (SAIL)

If you find this repository useful, please consider giving us a star :star: and citation:

@article{zhao2025workflows,
  title={On the Workflows and Smells of Leaderboard Operations (LBOps): An Exploratory Study of Foundation Model Leaderboards},
  author={Zhao, Zhimin and Bangash, Abdul Ali and C{\^o}go, Filipe Roseiro and Adams, Bram and Hassan, Ahmed E},
  journal={IEEE Transactions on Software Engineering},
  year={2025},
  publisher={IEEE}
}

If you want to contribute to this list (please do), welcome to propose a pull request.

If you have any suggestions, critiques, or questions regarding this list, welcome to raise issue.

Also, a leaderboard should be included if only:

  • It is actively maintained.
  • It is related to AI.

Table of Contents

Tools

NameDescription
Demo Leaderboard BackendDemo leaderboard backend helps users manage the leaderboard and handle submission requests, check this for details.
Evaluation Results on the HubEvaluation Results on the Hub enables model authors to store and display evaluation results in model cards by embedding structured metadata, making benchmark scores publicly accessible and comparable across models.
Kaggle Competition CreationKaggle Competition Creation enables you to design and launch custom competitions, leveraging your datasets to engage the data science community.

Challenges

NameDescription
AIcrowdAIcrowd hosts machine learning challenges and competitions across domains such as computer vision, NLP, and reinforcement learning, aimed at both researchers and practitioners.
AI HubAI Hub offers a variety of competitions to encourage AI solutions to real-world problems, with a focus on innovation and collaboration.
AI StudioAI Studio offers AI competitions mainly for computer vision, NLP, and other data-driven tasks, allowing users to develop and showcase their AI skills.
Allen Institute for AIThe Allen Institute for AI provides leaderboards and benchmarks on tasks in natural language understanding, commonsense reasoning, and other areas in AI research.
CodabenchCodabench is an open-source platform for benchmarking AI models, enabling customizable, user-driven challenges across various AI domains.
DataFountainDataFountain is a Chinese AI competition platform featuring challenges in finance, healthcare, and smart cities, encouraging solutions for industry-related problems.
DrivenDataDrivenData hosts machine learning challenges with a social impact, aiming to solve issues in areas, such as public health, disaster relief, and sustainable development.
DynabenchDynabench offers dynamic benchmarks where models are evaluated continuously, often involving human interaction, to ensure robustness in evolving AI tasks.
Eval AIEvalAI is a platform for hosting and participating in AI challenges, widely used by researchers for benchmarking models in tasks, such as image classification, NLP, and reinforcement learning.
Grand ChallengeGrand Challenge provides a platform for medical imaging challenges, supporting advancements in medical AI, particularly in areas, such as radiology and pathology.
HiltiHilti hosts challenges aimed at advancing AI and machine learning in the construction industry, with a focus on practical, industry-relevant applications.
InsightFaceInsightFace focuses on AI challenges related to face recognition, verification, and analysis, supporting advancements in identity verification and security.
KaggleKaggle is one of the largest platforms for data science and machine learning competitions, covering a broad range of topics from image classification to NLP and predictive modeling.
nuScenesnuScenes enables researchers to study challenging urban driving situations using the full sensor suite of a real self-driving car, facilitating research in autonomous driving.
Robust Reading CompetitionRobust Reading refers to the research area on interpreting written communication in unconstrained settings, with competitions focused on text recognition in real-world environments.
TianchiTianchi, hosted by Alibaba, offers a range of AI competitions, particularly popular in Asia, with a focus on commerce, healthcare, and logistics.

Rankings

Model Ranking

Comprehensive

NameDescription
AI Benchmarking HubAI Benchmarking Hub tracks and compares AI model performance in reasoning, coding, and knowledge tasks.
ArenaArena operates a chatbot arena where various foundation models compete based on user preferences across multiple categories: text generation, web development, computer vision, text-to-image synthesis, search capabilities, and coding assistance.
BenchGeckoBenchGecko is a comprehensive leaderboard that tracks thousands of models across 128 benchmarks, featuring cross-provider pricing comparisons, AI economy insights, an agent leaderboard, and an MCP server directory.
CompassRankCompassRank is a platform to offer a comprehensive, objective, and neutral evaluation reference of foundation models for the industry and research.
EvoClawEvoClaw is a leaderboard for evaluating and ranking AI agents across benchmark tasks.
FlagEvalFlagEval is a comprehensive platform for evaluating foundation models.
Generative AI LeaderboardsGenerative AI Leaderboard ranks the top-performing generative AI models based on various metrics.
Holistic Agent LeaderboardHAL is a standardized, cost-aware, and third-party leaderboard for evaluating agents.
Holistic Evaluation of Language ModelsHolistic Evaluation of Language Models (HELM) is a reproducible and transparent framework for evaluating foundation models.
HumanlayaHumanlaya is a comprehensive leaderboard for evaluating and comparing AI models across benchmarks.
InferenceBench.aiInferenceBench.ai is a benchmark for evaluating autonomous AI agents to optimize LLM inference workloads under fixed compute constraints.
Job BenchJob Bench is a benchmark leaderboard for evaluating AI models on job-related tasks.
KeygateKeygate is an AI model evaluation and comparison platform covering text, image, video, and speech models with rankings, side-by-side comparisons, and public methodology.
LLM StatsLLM Stats, the most comprehensive LLM leaderboard, benchmarks and compares API models using daily-updated, open-source community data on capability, price, speed, and context length.
LLM-Perf LeaderboardLLM-Perf Leaderboard aims to benchmark the performance of LLMs with different hardware, backends, and optimizations.
LLMPerfLLMPerf is a tool to evaluate the performance of LLMs using both load and correctness tests.
MSNP LeaderboardMSNP Leaderboard tracks and evaluates quantized GGUF models' performance on various GPU and CPU combinations using single-node setups via Ollama.
oobaboogaOobabooga is a benchmark to perform repeatable performance tests of LLMs with the oobabooga web UI.
PinchBenchPinchBench is a benchmark for evaluating and comparing AI agents in the OpenClaw environment across diverse tasks.
Scale LabsScale Labs is an AI research organization that evaluates and ranks AI models across various capability benchmarks.
SEAL LLM LeaderboardsSEAL LLM Leaderboards evaluate agentic, frontier, safety, and public-sentiment of the latest LLMs.
SEAL ShowdownSEAL Showdown ranks AI models based on how they perform in real-world use. Votes are blind, optional, and organic, so rankings reflect authentic preferences.
SkillsBenchSkillsBench is a leaderboard for evaluating and comparing the skills and capabilities of AI models.
SuperCLUESuperCLUE is a series of benchmarks for evaluating Chinese foundation models.
大模型修仙榜大模型修仙榜 (Large Model Cultivation Leaderboard) continuously tracks global model performance by combining authoritative evaluations such as Chatbot Arena, Artificial Analysis, SWE-bench, and Terminal-Bench across capability, coding, agent tasks, open-source deployment, stability, and cost-effectiveness.
Vals AIVal AI builds custom, industry-specific benchmarks using private datasets to provide unbiased third-party evaluations of LLM performance.
Vellum LLM LeaderboardVellum LLM Leaderboard shows a comparison of capabilities, price and context window for leading commercial and open-source LLMs.
XBenchXBench is a platform for evaluating and comparing AI models across diverse benchmarks and tasks.
Yupp LeaderboardYupp is a platform that enables users to compare outputs from multiple AI models side by side, select their preferred response, and provide feedback.

Text

NameDescription
ACLUEACLUE is an evaluation benchmark for ancient Chinese language comprehension.
African Languages LLM Eval LeaderboardAfrican Languages LLM Eval Leaderboard tracks progress and ranks performance of LLMs on African languages.
AGIEvalAGIEval is a human-centric benchmark to evaluate the general abilities of foundation models in tasks pertinent to human cognition and problem-solving.
AIR-BenchAIR-Bench is a benchmark to evaluate heterogeneous information retrieval capabilities of language models.
AI Energy Score LeaderboardAI Energy Score Leaderboard tracks and compares different models in energy efficiency.
AlignBenchAlignBench is a multi-dimensional benchmark for evaluating LLMs' alignment in Chinese.
AlpacaEvalAlpacaEval is an automatic evaluator designed for instruction-following LLMs.
ANGOANGO is a generation-oriented Chinese language model evaluation benchmark.
Arabic Tokenizers LeaderboardArabic Tokenizers Leaderboard compares the efficiency of LLMs in parsing Arabic in its different dialects and forms.
Arena-Hard-AutoArena-Hard-Auto is a benchmark for instruction-tuned LLMs.
AutoBench LLM LeaderboardAutoBench LLM Leaderboard is a platform where LLMs rank LLMs' responses in terms of performance, cost, and latency metrics.
AutoRaceAutoRace focuses on the direct evaluation of LLM reasoning chains with metric AutoRace (Automated Reasoning Chain Evaluation).
Auto ArenaAuto Arena is a benchmark in which various language model agents engage in peer-battles to evaluate their performance.
Auto-JAuto-J hosts evaluation results on the pairwise response comparison and critique generation tasks.
BABILongBABILong is a benchmark for evaluating the performance of language models in processing arbitrarily long documents with distributed facts.
BBLBBL (BIG-bench Lite) is a small subset of 24 diverse JSON tasks from BIG-bench. It is designed to provide a canonical measure of model performance, while being far cheaper to evaluate than the full set of more than 200 programmatic and JSON tasks in BIG-bench.
BeHonestBeHonest is a benchmark to evaluate honesty - awareness of knowledge boundaries (self-knowledge), avoidance of deceit (non-deceptiveness), and consistency in responses (consistency) - in LLMs.
BenBenchBenBench is a benchmark to evaluate the extent to which LLMs conduct verbatim training on the training set of a benchmark over the test set to enhance capabilities.
BenCzechMarkBenCzechMark (BCM) is a multitask and multimetric Czech language benchmark for LLMs with a unique scoring system that utilizes the theory of statistical significance.
BiGGen-BenchBiGGen-Bench is a comprehensive benchmark to evaluate LLMs across a wide variety of tasks.
BotChatBotChat is a benchmark to evaluate the multi-round chatting capabilities of LLMs through a proxy task.
CaselawQACaselawQA is a benchmark comprising legal classification tasks derived from the Supreme Court and Songer Court of Appeals legal databases.
Ch3EfCh3Ef is a benchmark to evaluate alignment with human expectations using 1002 human-annotated samples across 12 domains and 46 tasks based on the hhh principle.
Chain-of-Thought HubChain-of-Thought Hub is a benchmark to evaluate the reasoning capabilities of LLMs.
ChemBenchChemBench is a benchmark to evaluate the chemical knowledge and reasoning abilities of LLMs.
Chinese SimpleQAChinese SimpleQA is a Chinese benchmark to evaluate the factuality ability of language models to answer short questions.
CLBenchCLBench is a benchmark to evaluate the ability of LMs to learn new knowledge from provided context, covering domain knowledge reasoning, rule system application, procedural task execution, and empirical discovery.
CLEMCLEM is a framework designed for the systematic evaluation of chat-optimized LLMs as conversational agents.
CLEVACLEVA is a benchmark to evaluate LLMs on 31 tasks using 370K Chinese queries from 84 diverse datasets and 9 metrics.
Chinese Large Model LeaderboardChinese Large Model Leaderboard is a platform to evaluate the performance of Chinese LLMs.
CMMLUCMMLU is a benchmark to evaluate the performance of LLMs in various subjects within the Chinese cultural context.
CMMMUCMMMU is a benchmark to evaluate LMMs on tasks demanding college-level subject knowledge and deliberate reasoning in a Chinese context.
CommonGenCommonGen is a benchmark to evaluate generative commonsense reasoning by testing machines on their ability to compose coherent sentences using a given set of common concepts.
CompMixCompMix is a benchmark for heterogeneous question answering.
Compression Rate LeaderboardCompression Rate Leaderboard aims to evaluate tokenizer performance on different languages.
Compression LeaderboardCompression Leaderboard is a platform to evaluate the compression performance of LLMs.
CopyBenchCopyBench is a benchmark to evaluate the copying behavior and utility of language models as well as the effectiveness of methods to mitigate copyright risks.
CoTaEvalCoTaEval is a benchmark to evaluate the feasibility and side effects of copyright takedown methods for LLMs.
ConvReConvRe is a benchmark to evaluate LLMs' ability to comprehend converse relations.
CriticEvalCriticEval is a benchmark to evaluate LLMs' ability to make critique responses.
CS-BenchCS-Bench is a bilingual benchmark designed to evaluate LLMs' performance across 26 computer science subfields, focusing on knowledge and reasoning.
CUTECUTE is a benchmark to test the orthographic knowledge of LLMs.
CyberMetricCyberMetric is a benchmark to evaluate the cybersecurity knowledge of LLMs.
CzechBenchCzechBench is a benchmark to evaluate Czech language models.
C-EvalC-Eval is a Chinese evaluation suite for LLMs.
Decentralized ArenaDecentralized Arena hosts a decentralized and democratic platform for LLM evaluation, automating and scaling assessments across diverse, user-defined dimensions, including mathematics, logic, and science.
DecodingTrustDecodingTrust is a platform to evaluate the trustworthiness of LLMs.
Domain LLM LeaderboardDomain LLM Leaderboard is a platform to evaluate the popularity of domain-specific LLMs.
Enterprise Scenarios leaderboardEnterprise Scenarios Leaderboard tracks and evaluates the performance of LLMs on real-world enterprise use cases.
EQ-BenchEQ-Bench is a benchmark to evaluate aspects of emotional intelligence in LLMs.
European LLM LeaderboardEuropean LLM Leaderboard tracks and compares performance of LLMs in European languages.
EvalGPT.aiEvalGPT.ai hosts a chatbot arena to compare and rank the performance of LLMs.
Eval ArenaEval Arena measures noise levels, model quality, and benchmark quality by comparing model pairs across several LLM evaluation benchmarks with example-level analysis and pairwise comparisons.
Factuality LeaderboardFactuality Leaderboard compares the factual capabilities of LLMs.
FACTS GroundingFACTS Grounding is a benchmark to evaluate the ability of language models to ground their responses in factual information.
FanOutQAFanOutQA is a high quality, multi-hop, multi-document benchmark for LLMs using English Wikipedia as its knowledge base.
FastEvalFastEval is a toolkit for quickly evaluating instruction-following and chat language models on various benchmarks with fast inference and detailed performance insights.
FELMFELM is a meta benchmark to evaluate factuality evaluation benchmark for LLMs.
Fine-tuning LeaderboardFine-tuning Leaderboard is a platform to rank and showcase models that have been fine-tuned using open-source datasets or frameworks.
FollowBenchFollowBench is a multi-level fine-grained constraints following benchmark to evaluate the instruction-following capability of LLMs.
Forbidden Question DatasetForbidden Question Dataset is a benchmark containing 160 questions from 160 violated categories, with corresponding targets for evaluating jailbreak methods.
FuseReviewsFuseReviews aims to advance grounded text generation tasks, including long-form question-answering and summarization.
GAIAGAIA aims to test fundamental abilities that an AI assistant should possess.
GAVIEGAVIE is a GPT-4-assisted benchmark for evaluating hallucination in LMMs by scoring accuracy and relevancy without relying on human-annotated groundtruth.
GPT-FathomGPT-Fathom is an LLM evaluation suite, benchmarking 10+ leading LLMs as well as OpenAI's legacy models on 20+ curated benchmarks across 7 capability categories, all under aligned settings.
GrailQAStrongly Generalizable Question Answering (GrailQA) is a large-scale, high-quality benchmark for question answering on knowledge bases (KBQA) on Freebase with 64,331 questions annotated with both answers and corresponding logical forms in different syntax (i.e., SPARQL, S-expression, etc.).
Guerra LLM AI LeaderboardGuerra LLM AI Leaderboard compares and ranks the performance of LLMs across quality, price, performance, context window, and others.
Hallucinations LeaderboardHallucinations Leaderboard aims to track, rank and evaluate hallucinations in LLMs.
HalluQAHalluQA is a benchmark to evaluate the phenomenon of hallucinations in Chinese LLMs.
Hebrew LLM LeaderboardHebrew LLM Leaderboard tracks and ranks language models according to their success on various tasks on Hebrew.
HellaSwagHellaSwag is a benchmark to evaluate common-sense reasoning in LLMs.
Hughes Hallucination Evaluation Model leaderboardHughes Hallucination Evaluation Model leaderboard is a platform to evaluate how often a language model introduces hallucinations when summarizing a document.
Icelandic LLM leaderboardIcelandic LLM leaderboard tracks and compare models on Icelandic-language tasks.
IFEvalIFEval is a benchmark to evaluate LLMs' instruction following capabilities with verifiable instructions.
IL-TURIL-TUR is a benchmark for evaluating language models on monolingual and multilingual tasks focused on understanding and reasoning over Indian legal documents.
Indic LLM LeaderboardIndic LLM Leaderboard is platform to track and compare the performance of Indic LLMs.
Indico LLM LeaderboardIndico LLM Leaderboard evaluates and compares the accuracy of various language models across providers, datasets, and capabilities like text classification, key information extraction, and generative summarization.
InstructEvalInstructEval is a suite to evaluate instruction selection methods in the context of LLMs.
Italian LLM-LeaderboardItalian LLM-Leaderboard tracks and compares LLMs in Italian-language tasks.
JailbreakBenchJailbreakBench is a benchmark for evaluating LLM vulnerabilities through adversarial prompts.
Japanese Chatbot ArenaJapanese Chatbot Arena hosts the chatbot arena, where various LLMs compete based on their performance in Japanese.
Japanese LLM Roleplay BenchmarkJapanese LLM Roleplay Benchmark is a benchmark to evaluate the performance of Japanese LLMs in character roleplay.
JMMMUJMMMU (Japanese MMMU) is a multimodal benchmark to evaluate LMM performance in Japanese.
JustEvalJustEval is a powerful tool designed for fine-grained evaluation of LLMs.
KI-Benchmark-DeutschKI-Benchmark-Deutsch tracks and ranks the performance of LLMs on German-language tasks, including administrative, legal, and business German.
KoLAKoLA is a benchmark to evaluate the world knowledge of LLMs.
LaMPLaMP (Language Models Personalization) is a benchmark to evaluate personalization capabilities of language models.
Language Model CouncilLanguage Model Council (LMC) is a benchmark to evaluate tasks that are highly subjective and often lack majoritarian human agreement.
LawBenchLawBench is a benchmark to evaluate the legal capabilities of LLMs.
La LeaderboardLa Leaderboard evaluates and tracks LLM memorization, reasoning and linguistic capabilities in Spain, LATAM and Caribbean.
LogicKorLogicKor is a benchmark to evaluate the multidisciplinary thinking capabilities of Korean LLMs.
LongICLLongICLBench is a benchmark to evaluate long in-context learning evaluations for LLMs.
LooGLELooGLE is a benchmark to evaluate long context understanding capabilties of LLMs.
LAiWLAiW is a benchmark to evaluate Chinese legal language understanding and reasoning.
LLM Benchmarker SuiteLLM Benchmarker Suite is a benchmark to evaluate the comprehensive capabilities of LLMs.
Large Language Model Assessment in English ContextsLarge Language Model Assessment in English Contexts is a platform to evaluate LLMs in the English context.
Large Language Model Assessment in the Chinese ContextLarge Language Model Assessment in the Chinese Context is a platform to evaluate LLMs in the Chinese context.
LegalBenchLegalBench is a comprehensive benchmark for evaluating legal reasoning in LLMs.
LIBRALIBRA is a benchmark for evaluating LLMs' capabilities in understanding and processing long Russian text.
LibrAI-Eval GenAI LeaderboardLibrAI-Eval GenAI Leaderboard focuses on the balance between the LLM’s capability and safety in English.
LiveBenchLiveBench is a benchmark for LLMs to minimize test set contamination and enable objective, automated evaluation across diverse, regularly updated tasks.
LLM Confabulation/Hallucination LeaderboardLLM Confabulation/Hallucination Leaderboard evaluates LLMs based on how often they produce non-existent answers (confabulations or hallucinations) in response to misleading questions based on provided text documents, making it particularly useful for RAG scenarios.
LLMEvalLLMEval is a benchmark to evaluate the quality of open-domain conversations with LLMs.
Llmeval-Gaokao2024-MathLlmeval-Gaokao2024-Math is a benchmark for evaluating LLMs on 2024 Gaokao-level math problems in Chinese.
LLMHallucination LeaderboardHallucinations Leaderboard evaluates LLMs based on an array of hallucination-related benchmarks.
LLMs Disease Risk Prediction LeaderboardLLMs Disease Risk Prediction Leaderboard is a platform to evaluate LLMs on disease risk prediction.
LLM LeaderboardLLM Leaderboard tracks and evaluates LLM providers, enabling selection of the optimal API and model for user needs.
LLM ObservatoryLLM Observatory is a benchmark that assesses and ranks LLMs based on their performance in avoiding social biases across categories like LGBTIQ+ orientation, age, gender, politics, race, religion, and xenophobia.
LLM Price LeaderboardLLM Price Leaderboard tracks and compares LLM costs based on one million tokens.
LLM-AggreFactLLM-AggreFact is a fact-checking benchmark that aggregates most up-to-date publicly available datasets on grounded factuality evaluation.
LLM-LeaderboardLLM-Leaderboard is a joint community effort to create one central leaderboard for LLMs.
LMExamQALMExamQA is a benchmarking framework where a language model acts as an examiner to generate questions and evaluate responses in a reference-free, automated manner for comprehensive, equitable assessment.
LongBenchLongBench is a benchmark for assessing the long context understanding capabilities of LLMs.
LoongLoong is a long-context benchmark for evaluating LLMs' multi-document QA abilities across financial, legal, and academic scenarios.
Low-bit Quantized Open LLM LeaderboardLow-bit Quantized Open LLM Leaderboard tracks and compares quantization LLMs with different quantization algorithms.
LV-EvalLV-Eval is a long-context benchmark with five length levels and advanced techniques for accurate evaluation of LLMs on single-hop and multi-hop QA tasks across bilingual datasets.
LucyEvalLucyEval offers a thorough assessment of LLMs' performance in various Chinese contexts.
L-EvalL-Eval is a Long Context Language Model (LCLM) evaluation benchmark to evaluate the performance of handling extensive context.
M3KEM3KE is a massive multi-level multi-subject knowledge evaluation benchmark to measure the knowledge acquired by Chinese LLMs.
MetaCritiqueMetaCritique is a judge that can evaluate human-written or LLMs-generated critique by generating critique.
MINTMINT is a benchmark to evaluate LLMs' ability to solve tasks with multi-turn interactions by using tools and leveraging natural language feedback.
Meta Open LLM leaderboardThe Meta Open LLM leaderboard serves as a central hub for consolidating data from various open LLM leaderboards into a single, user-friendly visualization page.
MIMIC Clinical Decision Making LeaderboardMIMIC Clinical Decision Making Leaderboard tracks and evaluates LLms in realistic clinical decision-making for abdominal pathologies.
MixEvalMixEval is a benchmark to evaluate LLMs via by strategically mixing off-the-shelf benchmarks.
ML.ENERGY LeaderboardML.ENERGY Leaderboard evaluates the energy consumption of LLMs.
MMLUMMLU is a benchmark to evaluate the performance of LLMs across a wide array of natural language understanding tasks.
MMLU-by-task LeaderboardMMLU-by-task Leaderboard provides a platform for evaluating and comparing various ML models across different language understanding tasks.
MMLU-ProMMLU-Pro is a more challenging version of MMLU to evaluate the reasoning capabilities of LLMs.
ModelScope LLM LeaderboardModelScope LLM Leaderboard is a platform to evaluate LLMs objectively and comprehensively.
Model Evaluation LeaderboardModel Evaluation Leaderboard tracks and evaluates text generation models based on their performance across various benchmarks using Mosaic Eval Gauntlet framework.
MSTEBMSTEB is a benchmark for measuring the performance of text embedding models in Spanish.
MT-Bench-101MT-Bench-101 is a fine-grained benchmark for evaluating LLMs in multi-turn dialogues.
MY Malay LLM LeaderboardMY Malay LLM Leaderboard aims to track, rank, and evaluate open LLMs on Malay tasks.
NoChaNoCha is a benchmark to evaluate how well long-context language models can verify claims written about fictional books.
NPHardEvalNPHardEval is a benchmark to evaluate the reasoning abilities of LLMs through the lens of computational complexity classes.
Occiglot Euro LLM LeaderboardOcciglot Euro LLM Leaderboard compares LLMs in four main languages from the Okapi benchmark and Belebele (French, Italian, German, Spanish and Dutch).
OlympiadBenchOlympiadBench is a bilingual multimodal scientific benchmark featuring 8,476 Olympiad-level mathematics and physics problems with expert-level step-by-step reasoning annotations.
OlympicArenaOlympicArena is a benchmark to evaluate the advanced capabilities of LLMs across a broad spectrum of Olympic-level challenges.
OpenEvalOpenEval is a platform assessto evaluate Chinese LLMs.
OpenLLM Turkish leaderboardOpenLLM Turkish leaderboard tracks progress and ranks the performance of LLMs in Turkish.
Openness LeaderboardOpenness Leaderboard tracks and evaluates models' transparency in terms of open access to weights, data, and licenses, exposing models that fall short of openness standards.
Openness LeaderboardOpenness Leaderboard is a tool that tracks the openness of instruction-tuned LLMs, evaluating their transparency, data, and model availability.
OpenResearcherOpenResearcher contains the benchmarking results on various RAG-related systems as a leaderboard.
Open Arabic LLM LeaderboardOpen Arabic LLM Leaderboard tracks progress and ranks the performance of LLMs in Arabic.
Open Chinese LLM LeaderboardOpen Chinese LLM Leaderboard aims to track, rank, and evaluate open Chinese LLMs.
Open CoT LeaderboardOpen CoT Leaderboard tracks LLMs' abilities to generate effective chain-of-thought reasoning traces.
Open Dutch LLM Evaluation LeaderboardOpen Dutch LLM Evaluation Leaderboard tracks progress and ranks the performance of LLMs in Dutch.
Open ITA LLM LeaderboardOpen ITA LLM Leaderboard tracks progress and ranks the performance of LLMs in Italian.
Open Ko-LLM LeaderboardOpen Ko-LLM Leaderboard tracks progress and ranks the performance of LLMs in Korean.
Open LLM LeaderboardOpen LLM Leaderboard tracks progress and ranks the performance of LLMs in English.
Open MLLM LeaderboardOpen MLLM Leaderboard aims to track, rank and evaluate LLMs and chatbots.
Open MOE LLM LeaderboardOPEN MOE LLM Leaderboard assesses the performance and efficiency of various Mixture of Experts (MoE) LLMs.
Open Multilingual LLM Evaluation LeaderboardOpen Multilingual LLM Evaluation Leaderboard tracks progress and ranks the performance of LLMs in multiple languages.
Open PL LLM LeaderboardOpen PL LLM Leaderboard is a platform for assessing the performance of various LLMs in Polish.
Open Portuguese LLM LeaderboardOpen PT LLM Leaderboard aims to evaluate and compare LLMs in the Portuguese-language tasks.
Open Taiwan LLM leaderboardOpen Taiwan LLM leaderboard showcases the performance of LLMs on various Taiwanese Mandarin language understanding tasks.
Open-LLM-LeaderboardOpen-LLM-Leaderboard evaluates LLMs in language understanding and reasoning by transitioning from multiple-choice questions (MCQs) to open-style questions.
OPUS-MT DashboardOPUS-MT Dashboard is a platform to track and compare machine translation models across multiple language pairs and metrics.
ParsBenchParsBench provides toolkits for benchmarking LLMs based on the Persian language.
Persian LLM LeaderboardPersian LLM Leaderboard provides a reliable evaluation of LLMs in Persian Language.
Pinocchio ITA leaderboardPinocchio ITA leaderboard tracks and evaluates LLMs in Italian Language.
PL-MTEBPL-MTEB (Polish Massive Text Embedding Benchmark) is a benchmark for evaluating text embeddings in Polish across 28 NLP tasks.
PromptBenchPromptBench is a benchmark to evaluate the robustness of LLMs on adversarial prompts.
QAConvQAConv is a benchmark for question answering using complex, domain-specific, and asynchronous conversations as the knowledge source.
QuALITYQuALITY is a benchmark for evaluating multiple-choice question-answering with a long context.
RABBITSRABBITS is a benchmark to evaluate the robustness of LLMs by evaluating their handling of synonyms, specifically brand and generic drug names.
RakudaRakuda is a benchmark to evaluate LLMs based on how well they answer a set of open-ended questions about Japanese topics.
RedTeam ArenaRedTeam Arena is a red-teaming platform for LLMs.
Red Teaming Resistance BenchmarkRed Teaming Resistance Benchmark is a benchmark to evaluate the robustness of LLMs against red teaming prompts.
REFUTEREFUTE ranks models on scientific critique and epistemic calibration (judge-free): critique skill ≠ calibrated honesty on recent science summaries. Site: https://bgpt.pro/refute
ReST-MCTS*ReST-MCTS* is a reinforced self-training method that uses tree search and process reward inference to collect high-quality reasoning traces for training policy and reward models without manual step annotations.
Reviewer ArenaReviewer Arena hosts the reviewer arena, where various LLMs compete based on their performance in critiquing academic papers.
RoleEvalRoleEval is a bilingual benchmark to evaluate the memorization, utilization, and reasoning capabilities of role knowledge of LLMs.
RTEBRTEB (ReTrieval Embedding Benchmark) is a comprehensive benchmark for evaluating text retrieval models across multiple specialized domains including legal, finance, code, and healthcare.
Russian Chatbot ArenaChatbot Arena hosts a chatbot arena where various LLMs compete in Russian based on user satisfaction.
Russian SuperGLUERussian SuperGLUE is a benchmark for Russian language models, focusing on logic, commonsense, and reasoning tasks.
ScandEvalScandEval is a benchmark to evaluate LLMs on tasks in Scandinavian languages as well as German, Dutch, and English.
SCBenchSCBench is a benchmark to evaluate the long-context capabilities of LLMs across diverse tasks requiring shared context processing.
Science LeaderboardScience Leaderboard is a platform to evaluate LLMs' capabilities to solve science problems.
SciGLMSciGLM is a suite of scientific language models that use a self-reflective instruction annotation framework to enhance scientific reasoning by generating and revising step-by-step solutions to unlabelled questions.
SciKnowEvalSciKnowEval is a benchmark to evaluate LLMs based on their proficiency in studying extensively, enquiring earnestly, thinking profoundly, discerning clearly, and practicing assiduously.
SCROLLSSCROLLS is a benchmark to evaluate the reasoning capabilities of LLMs over long texts.
SeaExamSeaExam is a benchmark to evaluate LLMs for Southeast Asian (SEA) languages.
SeaEvalSeaEval is a benchmark to evaluate the performance of multilingual LLMs in understanding and reasoning with natural language, as well as comprehending cultural practices, nuances, and values.
SEA HELMSEA HELM is a benchmark to evaluate LLMs' performance across English and Southeast Asian tasks, focusing on chat, instruction-following, and linguistic capabilities.
SecEvalSecEval is a benchmark to evaluate cybersecurity knowledge of foundation models.
Self-Improving LeaderboardSelf-Improving Leaderboard (SIL) is a dynamic platform that continuously updates test datasets and rankings to provide real-time performance insights for open-source LLMs and chatbots.
SenseBenchSenseBench is a benchmark to evaluate how well LLMs disambiguate English words in context, maintained by Glite on its lexEN dataset of 4,861 polysemous items with WordNet 3.0 candidate senses, where every submitted run is re-verified in CI from the stored raw API responses.
SimpleBenchSimpleBench is a multiple-choice text benchmark where high school-level humans outperform all tested frontier LLMs, featuring 200+ questions on spatio-temporal reasoning, social intelligence, and linguistic adversarial robustness to test basic reasoning beyond memorized knowledge.
Spec-BenchSpec-Bench is a benchmark to evaluate speculative decoding methods across diverse scenarios.
StructEvalStructEval is a benchmark to evaluate LLMs by conducting structured assessments across multiple cognitive levels and critical concepts.
StructEval (Structural Outputs)StructEval (Structural Outputs) is a TMLR benchmark and leaderboard for evaluating LLM generation and conversion across 2,035 examples, 18 text and visual structured-output formats, and 44 task types.
Subquadratic LLM LeaderboardSubquadratic LLM Leaderboard evaluates LLMs with subquadratic/attention-free architectures (i.e. RWKV & Mamba).
SuperBenchSuperBench is a comprehensive system of tasks and dimensions to evaluate the overall capabilities of LLMs.
SuperGLUESuperGLUE is a benchmark to evaluate the performance of LLMs on a set of challenging language understanding tasks.
SuperLimSuperLim is a benchmark to evaluate the language understanding capabilities of LLMs in Swedish.
Swahili LLM-LeaderboardSwahili LLM-Leaderboard is a joint community effort to create one central leaderboard for LLMs.
TableQAEvalTableQAEval is a benchmark to evaluate LLM performance in modeling long tables and comprehension capabilities, such as numerical and multi-hop reasoning.
TAT-DQATAT-DQA is a benchmark to evaluate LLMs on the discrete reasoning over documents that combine both structured and unstructured information.
TAT-QATAT-QA is a benchmark to evaluate LLMs on the discrete reasoning over documents that combines both tabular and textual content.
Thai LLM LeaderboardThai LLM Leaderboard aims to track and evaluate LLMs in the Thai-language tasks.
The PileThe Pile is a benchmark to evaluate the world knowledge and reasoning ability of LLMs.
TOFUTOFU is a benchmark to evaluate the unlearning performance of LLMs in realistic scenarios.
Toloka LLM LeaderboardToloka LLM Leaderboard is a benchmark to evaluate LLMs based on authentic user prompts and expert human evaluation.
ToolbenchToolBench is a platform for training, serving, and evaluating LLMs specifically for tool learning.
Toxicity LeaderboardToxicity Leaderboard evaluates the toxicity of LLMs.
Tracking AI leaderboardTracking AI leaderboard is an online ranking platform of AI models based on their performance on various IQ-style evaluations.
Trustbit LLM LeaderboardsTrustbit LLM Leaderboards is a platform that provides benchmarks for building and shipping products with LLMs.
TrustLLMTrustLLM is a benchmark to evaluate the trustworthiness of LLMs.
TuringAdviceTuringAdvice is a benchmark for evaluating language models' ability to generate helpful advice for real-life, open-ended situations.
TutorEvalTutorEval is a question-answering benchmark which evaluates how well an LLM tutor can help a user understand a chapter from a science textbook.
T-EvalT-Eval is a benchmark for evaluating the tool utilization capability of LLMs.
UGI LeaderboardUGI Leaderboard measures and compares the uncensored and controversial information known by LLMs.
UltraEvalUltraEval is an open-source framework for transparent and reproducible benchmarking of LLMs across various performance dimensions.
VCRVisual Commonsense Reasoning (VCR) is a benchmark for cognition-level visual understanding, requiring models to answer visual questions and provide rationales for their answers.
ViDoReViDoRe is a benchmark to evaluate retrieval models on their capacity to match queries to relevant documents at the page level.
VLLMs LeaderboardVLLMs Leaderboard aims to track, rank and evaluate open LLMs and chatbots.
VMLUVMLU is a benchmark to evaluate overall capabilities of foundation models in Vietnamese.
Vinayak Multistep Recursive Reasoning Benchmark (VMRRB)VMRRB is a benchmark for evaluating advanced reasoning, recursive dependency resolution, and robustness in dynamic, noisy, and structurally challenging environments. GitHub.
WildBenchWildBench is a benchmark for evaluating language models on challenging tasks that closely resemble real-world applications.
XiezhiXiezhi is a benchmark for holistic domain knowledge evaluation of LLMs.
Yanolja ArenaYanolja Arena host a model arena to evaluate the capabilities of LLMs in summarizing and translating text.
Yet Another LLM LeaderboardYet Another LLM Leaderboard is a platform for tracking, ranking, and evaluating open LLMs and chatbots.
ZebraLogicZebraLogic is a benchmark evaluating LLMs' logical reasoning using Logic Grid Puzzles, a type of Constraint Satisfaction Problem (CSP).
ZeroSumEvalZeroSumEval is a competitive evaluation framework for LLMs using multiplayer simulations with clear win conditions.

Code

NameDescription
Aider LLM LeaderboardsAider LLM Leaderboards evaluate LLM's ability to follow system prompts to edit code.
AndroidWorldAndroidWorld is a comprehensive testing framework to evaluate the abilities of AI models, specifically autonomous agents, in controlling and interacting with a mobile device.
APEX AI BenchmarksAPEX AI Benchmarks is a benchmark suite for evaluating AI model performance on coding and software engineering tasks.
AppWorldAppWorld is a high-fidelity execution environment of 9 day-to-day apps, operable via 457 APIs, populated with digital activities of ~100 people living in a simulated world.
Berkeley Function-Calling LeaderboardBerkeley Function-Calling Leaderboard evaluates the ability of LLMs to call functions (also known as tools) accurately.
BigCodeBenchBigCodeBench is a benchmark for code generation with practical and challenging programming tasks.
Big Code Models LeaderboardBig Code Models Leaderboard is a platform to track and evaluate the performance of LLMs on code-related tasks.
BIRDBIRD is a benchmark to evaluate the performance of text-to-SQL parsing systems.
CanAiCode LeaderboardCanAiCode Leaderboard is a platform to evaluate the code generation capabilities of LLMs.
ClassEvalClassEval is a benchmark to evaluate LLMs on class-level code generation.
CodeApexCodeApex is a benchmark to evaluate LLMs' programming comprehension through multiple-choice questions and code generation with C++ algorithm problems.
CodeScopeCodeScope is a benchmark to evaluate LLM coding capabilities across 43 languages and 8 tasks, considering difficulty, efficiency, and length.
CodeTransOceanCodeTransOcean is a benchmark to evaluate code translation across a wide variety of programming languages, including popular, niche, and LLM-translated code.
Code LinguaCode Lingua is a benchmark to compare the ability of code models to understand what the code implements in source languages and translate the same semantics in target languages.
CodeClashCodeClash benchmarks LLMs on goal-oriented software engineering through competitive coding tournaments.
CodeEloCodeElo is a benchmark to evaluate LLMs on code generation through competitive Elo-style ratings.
Coding LLMs LeaderboardCoding LLMs Leaderboard is a platform to evaluate and rank LLMs across various programming tasks.
Commit-0Commit-0 is a from-scratch AI coding challenge to rebuild 54 core Python libraries, ensuring they pass unit tests with significant test coverage, lint/type checking, and cloud-based distributed development.
CompileBenchCompileBench is a benchmark to evaluate how well LLMs can handle real-world software engineering challenges, including dependency hell, legacy toolchains, and cryptic compile errors.
CRUXEvalCRUXEval is a benchmark to evaluate code reasoning, understanding, and execution capabilities of LLMs.
CSpiderCSpider is a benchmark to evaluate systems' ability to generate SQL queries from Chinese natural language across diverse, complex, and cross-domain databases.
CyberSecEvalCyberSecEval is a benchmark to evaluate the cybersecurity of LLMs as coding assistants.
DevEvalDevEval is a code generation benchmark collected through a rigorous pipeline. DevEval contains 1,825 testing samples, collected from 115 real-world code repositories and covering 10 programming topics.
DevOps AI Assistant Open LeaderboardDevOps AI Assistant Open Leaderboard tracks, ranks, and evaluates DevOps AI Assistants across knowledge domains.
DevOps-EvalDevOps-Eval is a benchmark to evaluate code models in the DevOps/AIOps field.
DevOps-GymDevOps-Gym is a benchmark to evaluate AI agents across the complete software DevOps cycle, including build & configuration, monitoring, issue resolution, and test generation.
DomainEvalDomainEval is an auto-constructed benchmark for multi-domain code generation.
Dr.SpiderDr.Spider is a benchmark to evaluate the robustness of text-to-SQL models using different perturbation test sets.
EffiBenchEffiBench is a benchmark to evaluate the efficiency of LLMs in code generation.
EvalPlusEvalPlus is a benchmark to evaluate the code generation performance of LLMs.
EvoCodeBenchEvoCodeBench is an evolutionary code generation benchmark aligned with real-world code repositories.
EvoEvalEvoEval is a benchmark to evaluate the coding abilities of LLMs, created by evolving existing benchmarks into different targeted domains.
InfiBenchInfiBench is a benchmark to evaluate code models on answering freeform real-world code-related questions.
InterCodeInterCode is a benchmark to standardize and evaluate interactive coding with execution feedback.
Julia LLM LeaderboardJulia LLM Leaderboard is a platform to compare code models' abilities in generating syntactically correct Julia code, featuring structured tests and automated evaluations for easy and collaborative benchmarking.
LiveCodeBenchLiveCodeBench is a benchmark to evaluate code models across code-related scenarios over time.
LiveCodeBench GSOLiveCodeBench GSO is a benchmark to evaluate code models on generation, self-repair, and optimization tasks.
LiveCodeBench ProLiveCodeBench Pro evaluates LLMs on their ability to generate solutions for programming problems. The benchmark includes problems of varying difficulty levels from different competitive programming platforms.
Long Code ArenaLong Code Arena is a suite of benchmarks for code-related tasks with large contexts, up to a whole code repository.
McEvalMcEval is a massively multilingual code evaluation benchmark covering 40 languages (16K samples in 44 total), encompassing multilingual code generation, multilingual code explanation, and multilingual code completion tasks.
Memorization or Generation of Big Code Models LeaderboardMemorization or Generation of Big Code Models Leaderboard tracks and compares code generation models' performance.
MLE-benchMLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering
Multi-SWE-benchMulti-SWE-bench is a multi-lingual GitHub issue resolving benchmark for code agents.
NaturalCodeBenchNaturalCodeBench is a benchmark to mirror the complexity and variety of scenarios in real coding tasks.
Nexus Function Calling LeaderboardNexus Function Calling Leaderboard is a platform to evaluate code models on performing function calling and API usage.
NL2SQL360NL2SQL360 is a comprehensive evaluation framework for comparing and optimizing NL2SQL methods across various application scenarios.
OpsEvalOpsEval is a benchmark designed to assess the performance of LLMs in IT Operations (AIOps).
PECCPECC is a benchmark that evaluates code generation by requiring models to comprehend and extract problem requirements from narrative-based descriptions to produce syntactically accurate solutions.
ProgramBenchProgramBench is a benchmark to evaluate LLMs on programming tasks.
PyBenchPyBench is a benchmark evaluating LLM on real-world coding tasks including chart analysis, text analysis, image/ audio editing, complex math and software/website development.
RACERACE is a benchmark to evaluate the ability of LLMs to generate code that is correct and meets the requirements of real-world development scenarios.
RepairBenchRepairBench is a benchmark to evaluate the program repair capabilities of AI models by testing their ability to fix real-world software bugs.
ResearchCodeBenchResearchCodeBench is a benchmark for evaluating LLMs on their ability to translate novel machine learning research papers into executable code.
RepoQARepoQA is a benchmark to evaluate the long-context code understanding ability of LLMs.
RubberDuckBenchRubberDuckBench is a benchmark to evaluate LLMs on code explanation capabilities, inspired by the rubber duck debugging technique.
ScreenSpotScreenSpot is a benchmark to evaluate the ability of VLMs to perform GUI grounding across various platforms, including mobile (iOS, Android), desktop (macOS, Windows), and web environments, based on over 1,200 instructions.
ScreenSpot-ProScreenSpot-Pro is a benchmark to evaluate the ability of multi-modal large language models (MLLMs) to accurately locate specific GUI elements in complex, high-resolution desktop applications.
SciCodeSciCode is a benchmark designed to evaluate language models in generating code to solve realistic scientific research problems.
Software Engineering ArenaSoftware Engineering Arena is an open-source initiative to transparently evaluate and track AI assistants across real-world software engineering tasks.
SolidityBenchSolidityBench is a benchmark to evaluate and rank the ability of LLMs in generating and auditing smart contracts.
SpiderSpider is a benchmark to evaluate the performance of natural language interfaces for cross-domain databases.
SRE Skills BenchSRE Skills Bench is a benchmark to evaluate the site reliability engineering skills of LLMs.
StableToolBenchStableToolBench is a benchmark to evaluate tool learning that aims to provide a well-balanced combination of stability and reality.
SWE-benchSWE-bench is a benchmark for evaluating LLMs on real-world software issues collected from GitHub.
SWE-bench-LiveSWE-bench-Live is a live benchmark to evaluate an AI system's ability to complete real-world software engineering tasks.
SWE-Bench ProSWE-Bench Pro is a benchmark to evaluate an AI system's ability for challenging long-horizon software engineering tasks.
SWE-EffiSWE-Effi is an evaluation framework to evaluate the effectiveness of SWE agents by balancing performance metrics with resource consumption.
SWE-Model-ArenaSWE-Model-Arena provides a platform for software developers to compare the performance of different FMs on software engineering tasks.
SWE-rebenchSWE-rebench is a continuously updated benchmark designed to provide more accurate and reliable evaluations of software engineering LLMs by using real-world tasks from GitHub that are less prone to data contamination.
Terminal-BenchTerminal-Bench is a benchmark to measure the capabilities of AI agents in a terminal environment.
Turing Machine Programming BenchmarkMeasures LLMs' ability to solve algorithmic tasks by programming a Turing machine.
UI-I2E-BenchUI-I2E-Bench is a benchmark for GUI visual grounding. It incorporates implicit instructions and long-tail UI element types, with element-to-screen size ratios that better reflect real-world scenarios.
WebApp1KWebApp1K is a benchmark to evaluate LLMs on their abilities to develop real-world web applications.
WebDev ArenaWebDev Arena hosts a chatbot arena where various LLMs compete based on website development.
WeirdMLWeirdML is a benchmark to evaluate LLMs on their ability to solve novel and unusual machine learning tasks through iterative code generation, debugging, and optimization.
WILDSWILDS is a benchmark of in-the-wild distribution shifts spanning diverse data modalities and applications, from tumor identification to wildlife monitoring to poverty mapping.

Image

NameDescription
Abstract ImageAbstract Image is a benchmark to evaluate multimodal LLMs (MLLM) in understanding and visually reasoning about abstract images, such as maps, charts, and layouts.
AesBenchAesBench is a benchmark to evaluate MLLMs on image aesthetics perception.
ARC-AGI LeaderboardARC-AGI leaderboard publicly tracks and compares the efficiency and capability of artificial intelligence systems against human performance on the ARC-AGI-2 benchmark as part of an open-source research competition aimed at advancing artificial general intelligence.
BabyVisionBabyVision is a benchmark to evaluate the fundamental visual understanding capabilities of vision models.
BLINKBLINK is a benchmark to evaluate the core visual perception abilities of MLLMs.
BlinkCodeBlinkCode is a benchmark to evaluate MLLMs across 15 vision-language models (VLMs) and 9 tasks, measuring accuracy and image reconstruction performance.
Blueprint BenchBlueprint Bench is a benchmark to evaluate foundation models' capabilities in understanding, analyzing, and extracting information from architectural blueprints and technical diagrams.
ChartMimicChartMimic is a benchmark to evaluate the visually-grounded code generation capabilities of large multimodal models using charts and textual instructions.
CharXivCharXiv is a benchmark to evaluate chart understanding capabilities of MLLMs.
ConTextualConTextual is a benchmark to evaluate MLLMs across context-sensitive text-rich visual reasoning tasks.
CORE-MMCORE-MM is a benchmark to evaluate the open-ended visual question-answering (VQA) capabilities of MLLMs.
Design ArenaDesign Arena is a leaderboard to evaluate AI models for design-related tasks.
DreamBench++DreamBench++ is a human-aligned benchmark automated by multimodal models for personalized image generation.
EgoPlan-BenchEgoPlan-Bench is a benchmark to evaluate planning abilities of MLLMs in real-world, egocentric scenarios.
HallusionBenchHallusionBench is a benchmark to evaluate the image-context reasoning capabilities of MLLMs.
ImageBenchImageBench is a benchmark and comparison site for text-to-image models, with side-by-side outputs, prompt pass rates, and methodology pages.
InfiMM-EvalInfiMM-Eval is a benchmark to evaluate the open-ended VQA capabilities of MLLMs.
LRVSF LeaderboardLRVSF Leaderboard is a platform to evaluate LLMs regarding image similarity search in fashion.
LVLM LeaderboardLVLM Leaderboard is a platform to evaluate the visual reasoning capabilities of MLLMs.
M3CoTM3CoT is a benchmark for multi-domain multi-step multi-modal chain-of-thought of MLLMs.
MageBenchMageBench is a benchmark to evaluate multimodal large language models.
MementosMementos is a benchmark to evaluate the reasoning capabilities of MLLMs over image sequences.
MJ-BenchMJ-Bench is a benchmark to evaluate multimodal judges in providing feedback for image generation models across four key perspectives: alignment, safety, image quality, and bias.
MLLM-as-a-JudgeMLLM-as-a-Judge is a benchmark with human annotations to evaluate MLLMs' judging capabilities in scoring, pair comparison, and batch ranking tasks across multimodal domains.
MLLM-BenchMLLM-Bench is a benchmark to evaluate the visual reasoning capabilities of MLVMs.
MMBenchMMBench is a collection of benchmarks to evaluate the multi-modal reasoning capabilities of MLLMs.
MMEMME is a benchmark to evaluate the visual reasoning capabilities of MLLMs.
MME-RealWorldMME-RealWorld is a large-scale, high-resolution benchmark featuring 29,429 human-annotated QA pairs across 43 tasks.
MMIUMMIU (Ultimodal Multi-image Understanding) is a benchmark to evaluate MLLMs across 7 multi-image relationships, 52 tasks, 77K images, and 11K curated multiple-choice questions.
MMMUMMMU is a benchmark to evaluate the performance of multimodal models on tasks that demand college-level subject knowledge and expert-level reasoning across various disciplines.
MMRMMR is a benchmark to evaluate the robustness of MLLMs in visual understanding by assessing their ability to handle leading questions, rather than just accuracy in answering.
MMSearchMMSearch is a benchmark to evaluate the multimodal search performance of LMMs.
MMStarMMStar is a benchmark to evaluate the multi-modal capacities of MLLMs.
MMT-BenchMMT-Bench is a benchmark to evaluate MLLMs across a wide array of multimodal tasks that require expert knowledge as well as deliberate visual recognition, localization, reasoning, and planning.
MM-NIAHMM-NIAH (Needle In A Multimodal Haystack) is a benchmark to evaluate MLLMs' ability to comprehend long multimodal documents through retrieval, counting, and reasoning tasks involving both text and image data.
MTVQAMTVQA is a multilingual visual text comprehension benchmark to evaluate MLLMs.
Multimodal Hallucination LeaderboardMultimodal Hallucination Leaderboard compares MLLMs based on hallucination levels in various tasks.
MULTI-BenchmarkMULTI-Benchmark is a benchmark to evaluate MLLMs on understanding complex tables and images, and reasoning with long context.
NPHardEval4VNPHardEval4V is a benchmark to evaluate the reasoning abilities of MLLMs through the lens of computational complexity classes.
OCRBenchOCRBench is a benchmark to evaluate the OCR capabilities of multimodal models.
PCA-BenchPCA-Bench is a benchmark to evaluate the embodied decision-making capabilities of multimodal models.
Q-BenchQ-Bench is a benchmark to evaluate the visual reasoning capabilities of MLLMs.
RewardBenchRewardBench is a benchmark to evaluate the capabilities and safety of reward models.
ScienceQAScienceQA is a benchmark used to evaluate the multi-hop reasoning ability and interpretability of AI systems in the context of answering science questions.
SciGraphQASciGraphQA is a benchmark to evaluate the MLLMs in scientific graph question-answering.
SEED-BenchSEED-Bench is a benchmark to evaluate the text and image generation of multimodal models.
URIALURIAL is a benchmark to evaluate the capacity of language models for alignment without introducing the factors of fine-tuning (learning rate, data, etc.), which are hard to control for fair comparisons.
UPD LeaderboardUPD Leaderboard is a platform to evaluate the trustworthiness of MLLMs in unsolvable problem detection.
Vibe-EvalVibe-Eval is a benchmark to evaluate MLLMs for challenging cases.
VideoHallucerVideoHallucer is a benchmark to detect hallucinations in MLLMs.
VisIT-BenchVisIT-Bench is a benchmark to evaluate the instruction-following capabilities of MLLMs for real-world use.
Waymo Open Dataset ChallengesWaymo Open Dataset Challenges hold diverse self-driving datasets to evaluate ML models.
WHOOPS!WHOOPS! is a benchmark to evaluate the visual commonsense reasoning abilities of MLLMs.
WildVision-BenchWildVision-Bench is a benchmark to evaluate VLMs in the wild with human preferences.
WildVision ArenaWildVision Arena hosts the chatbot arena where various MLLMs compete based on their performance in visual understanding.
WorldVQAWorldVQA is a benchmark to evaluate visual question answering capabilities of MLLMs.

Video

NameDescription
ChronoMagic-BenchChronoMagic-Bench is a benchmark to evaluate video models' ability to generate time-lapse videos with high metamorphic amplitude and temporal coherence across physics, biology, and chemistry domains using free-form text control.
DREAM-1KDREAM-1K is a benchmark to evaluate video description performance on 1,000 diverse video clips featuring rich events, actions, and motions from movies, animations, stock videos, YouTube, and TikTok-style short videos.
LongVideoBenchLongVideoBench is a benchmark to evaluate the capabilities of video models in answering referred reasoning questions, which are dependent on long frame inputs and cannot be well-addressed by a single frame or a few sparse frames.
LVBenchLVBench is a benchmark to evaluate multimodal models on long video understanding tasks requiring extended memory and comprehension capabilities.
MLVUMLVU is a benchmark to evaluate video models in multi-task long video understanding.
MMOUMMOU (Massive Multi-Task Omni Understanding) is a benchmark to evaluate omni-modal models on long-form audio-visual reasoning over real-world web videos.
MMToM-QAMMToM-QA is a multimodal benchmark to evaluate machine Theory of Mind (ToM), the ability to understand people's minds.
MVBenchMVBench is a benchmark to evaluate the temporal understanding capabilities of video models in dynamic video tasks.
OpenVLM Video LeaderboardOpenVLM Video Leaderboard is a platform showcasing the evaluation results of 30 different VLMs on video understanding benchmarks using the VLMEvalKit framework.
TempCompassTempCompass is a benchmark to evaluate Video LLMs' temporal perception using 410 videos and 7,540 task instructions across 11 temporal aspects and 4 task types.
VBenchVBench is a benchmark to evaluate video generation capabilities of video models.
VideoNIAHVideoNIAH is a benchmark to evaluate the fine-grained understanding and spatio-temporal modeling capabilities of video models.
VideoPhyVideoPhy is a benchmark to evaluate generated videos for adherence to physical commonsense in real-world material interactions.
VideoScoreVideoScore is a benchmark to evaluate text-to-video generative models on five key dimensions.
VideoVistaVideoVista is a benchmark with 25,000 questions from 3,400 videos across 14 categories, covering 19 understanding and 8 reasoning tasks.
Video-BenchVideo-Bench is a benchmark to evaluate the video-exclusive understanding, prior knowledge incorporation, and video-based decision-making abilities of video models.
Video-MMEVideo-MME is a benchmark to evaluate the video analysis capabilities of video models.

Math

NameDescription
AbelAbel is a platform to evaluate the mathematical capabilities of LLMs.
FrontierMathFrontierMath is a benchmark to evaluate advanced mathematical reasoning capabilities of AI models.
ImProofBenchImProofBench is a benchmark to evaluate the mathematical proof generation and reasoning capabilities of LLMs.
MathArenaMathArena is a platform for evaluation of LLMs on the latest math competitions and olympiads.
MathBenchMathBench is a multi-level difficulty mathematics evaluation benchmark for LLMs.
MathEvalMathEval is a benchmark to evaluate the mathematical capabilities of LLMs.
MathUserEvalMathUserEval is a benchmark featuring university exam questions and math-related queries derived from simulated conversations with experienced annotators.
MathVerseMathVerse is a benchmark to evaluate vision-language models in interpreting and reasoning with visual information in mathematical problems.
MathVistaMathVista is a benchmark to evaluate mathematical reasoning in visual contexts.
MATH-VMATH-Vision (MATH-V) is a benchmark of 3,040 visually contextualized math problems from competitions, covering 16 disciplines and 5 difficulty levels to evaluate LMMs' mathematical reasoning.
Open Multilingual Reasoning LeaderboardOpen Multilingual Reasoning Leaderboard tracks and ranks the reasoning performance of LLMs on multilingual mathematical reasoning benchmarks.
PutnamBenchPutnamBench is a benchmark to evaluate the formal mathematical reasoning capabilities of LLMs on the Putnam Competition.
SciBenchSciBench is a benchmark to evaluate the reasoning capabilities of LLMs for solving complex scientific problems.
TabMWPTabMWP is a benchmark to evaluate LLMs in mathematical reasoning tasks that involve both textual and tabular data.
We-MathWe-Math is a benchmark to evaluate the human-like mathematical reasoning capabilities of LLMs with problem-solving principles beyond the end-to-end performance.

Agent

NameDescription
Agent ArenaAgent Arena ranks AI agents, models, tools, and frameworks using ELO-Style ratings from battles, offering insights into agent capabilities across various categories and leverages battle data to evaluate individual agent components.
AgentBenchAgentBench is the benchmark to evaluate language model-as-Agent across a diverse spectrum of different environments.
AgentBoardAgentBoard is a benchmark for multi-turn LLM agents, complemented by an analytical evaluation board for detailed model assessment beyond final success rates.
Agents' Last ExamAgents' Last Exam is a benchmark for evaluating generalist AI agents on long-horizon, economically valuable professional workflows.
AgentStudioAgentStudio is an integrated solution featuring in-depth benchmark suites, realistic environments, and comprehensive toolkits.
AssistantBenchAssistantBench aims to evaluate the ability of web agents to assist with real and time-consuming tasks.
BenchClawBenchClaw is a benchmark to evaluate AI agents on autonomous scientific paper writing and peer review using a multi-judge evaluation process.
BrowseCompBrowseComp is a benchmark to evaluate the ability of AI agents to locate hard-to-find information.
BrowserGym LeaderboardBrowserGym is a gym environment to evaluate LLMs, VLMs, and agents on web navigation tasks.
CharacterEvalCharacterEval is a benchmark to evaluate Role-Playing Conversational Agents (RPCAs) using multi-turn dialogues and character profiles, with metrics spanning four dimensions.
CEO-BenchCEO-Bench is a benchmark to evaluate AI agents on steering a simulated AI startup over a 500-day horizon.
ClawBenchClawBench is a browser-agent benchmark covering 283 everyday web tasks (V1 153 + V2 130) across 163 live production websites in 15 categories, with two-stage scoring (HTTP-request interception + LLM judge).
Claw-EvalClaw-Eval is a benchmark to evaluate LLM agents on real-world tasks, featuring 139 tasks across 15 services with Docker sandbox isolation and human-verified grading.
ClawWorkClawWork is a real-world economic benchmark where AI agents complete professional tasks spanning 44 occupations, earning income by performing quality work while managing token costs and maintaining economic solvency.
Computer Agent ArenaComputer Agent Arena is an open evaluation platform where users can compare LLM/VLM-based AI agents performing real-world computer tasks, ranging from general computer use to specialized workflows like coding, data analysis, and video editing.
CyberGymCyberGym is a benchmark to evaluate AI agents on cybersecurity tasks.
Galileo Agent LeaderboardGalileo Agent Leaderboard is an open evaluation platform to track and evaluate LLM agents in task completion and tool calling across business domains.
GTAGTA is a benchmark to evaluate the tool-use capability of LLM-based agents in real-world scenarios.
HlidoHlido ranks and reviews AI agents with Laddoo Scores, claim-vs-evidence tables, C2PA-signed proof artifacts, a public Hugging Face dataset, and a multi-tool MCP server.
HippoCampHippoCamp is a benchmark to evaluate contextual agents on their ability to search, perceive, and reason over realistic, multimodal personal file systems on personal computers.
Leetcode-Hard GymLeetcode-Hard Gym is an RL environment interface to LeetCode's submission server for evaluating codegen agents.
LLM Colosseum LeaderboardLLM Colosseum Leaderboard is a platform to evaluate LLMs by fighting in Street Fighter 3.
MAgICMAgIC is a benchmark to measure the abilities of cognition, adaptability, rationality and collaboration of LLMs within multi-agent systems.
MCP BenchMCP Bench is a benchmark to evaluate AI models on Model Context Protocol (MCP) server interactions and tool-use capabilities.
MCPMarkMCPMark is a benchmark to evaluate model and agent capabilities in real-world MCP tasks.
MCP UniverseMCP Universe is a leaderboard to compare AI model performance on MCP (Model Context Protocol) tasks.
Olas Predict BenchmarkOlas Predict Benchmark is a benchmark to evaluate agents on historical and future event forecasting.
OSWorldOSWorld is a benchmark to evaluate multimodal AI agents on their ability to perform 369 realistic, open-ended tasks within a virtual computer environment across various applications and operating systems.
OSWorld-MCPOSWorld-MCP is a benchmark to evaluate AI agents on real-world computer tasks using the Model Context Protocol (MCP).
SEC-benchSEC-bench is a benchmark of LLM agents on real-world software security tasks.
TravelPlannerTravelPlanner is a benchmark to evaluate LLM agents in tool use and complex planning within multiple constraints.
VABVisualAgentBench (VAB) is a benchmark to evaluate and develop LMMs as visual foundation agents, which comprises 5 distinct environments across 3 types of representative visual agent tasks.
VisualWebArenaVisualWebArena is a benchmark to evaluate the performance of multimodal web agents on realistic visually grounded tasks.
WebArenaWebArena is a standalone, self-hostable web environment to evaluate autonomous agents.
WildClawBenchWildClawBench is an in-the-wild benchmark to evaluate AI agents in the OpenClaw Environment.
YC-BenchYC-Bench is a long-horizon deterministic benchmark for LLM agents, where agents play CEO of an AI startup over a simulated 1–3 year run, managing decisions across prestige specialisation, employee allocation, cash flow, and deadline risk.
τ-Benchτ-bench is a benchmark that emulates dynamic conversations between a language model-simulated user and a language agent equipped with domain-specific API tools and policy guidelines.

Research

NameDescription
DeepResearch-Bench LeaderboardDeepResearch-Bench Leaderboard evaluates AI models on deep research tasks.

Business

NameDescription
AI-TraderAI-Trader is a fully autonomous trading benchmark to compare the performance of different AI models in trading NASDAQ 100 stocks.
Aiera LeaderboardAiera Leaderboard evaluates LLM performance on financial intelligence tasks, including speaker assignments, speaker change identification, abstractive summarizations, calculation-based Q&A, and financial sentiment tagging.
Alpha ArenaAlpha Arena is a benchmark to evaluate AI models' investing abilities.
BookSQLBookSQL is a benchmark to evaluate Text-to-SQL systems in the finance and accounting domain across various industries with a dataset of 1 million transactions from 27 businesses.
CFLUECFLUE is a benchmark to evaluate LLMs' understanding and processing capabilities in the Chinese fin

常见问题

What is awesome-ai-leaderboard?

awesome-ai-leaderboard is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by SAILResearch. A curated list of awesome leaderboard-oriented resources for AI domain. It has 379 GitHub stars.

Is awesome-ai-leaderboard safe to use?

Yes. awesome-ai-leaderboard passed SkillsLLM's automated security scan — a dependency vulnerability audit plus prompt-injection heuristics — with no high-severity issues. You can read the full report in the Security Report section on this page.

How do I install awesome-ai-leaderboard?

Clone the repository with "git clone https://github.com/SAILResearch/awesome-ai-leaderboard" and add it to your Claude Code skills directory (see the Installation section above).

Are there alternatives to awesome-ai-leaderboard?

Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh awesome-ai-leaderboard against similar tools.

评论 (0)

暂无评论,成为第一个分享想法的人!

ECC

by affaan-m

10

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

242,21936,702JavaScript
AI 智能体ai-agentsanthropicclaude-code
查看详情
15

An agentic skills framework & software development methodology that works.

234,96620,863Shell
AI 智能体ai-agentsbrainstorming
查看详情

hermes-agent

by NousResearch

10

The agent that grows with you

234,43747,175Python
AI 智能体ai-agentsagent-orchestration
查看详情

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

185,94028,768JavaScript
AI 智能体ai-agentsanthropicclaude-code
查看详情

cc-switch

by farion1231

3

A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io

128,8688,826Rust
AI 智能体claude-codeai-tools
查看详情

claude-code

by anthropics

Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.

120,03119,897Shell
AI 智能体
查看详情

开发者还喜欢

基于喜欢此 Skill 的开发者投票和收藏

ECC

by affaan-m

10

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

242,21936,702JavaScript
AI 智能体ai-agentsanthropicclaude-code
查看详情
15

An agentic skills framework & software development methodology that works.

234,96620,863Shell
AI 智能体ai-agentsbrainstorming
查看详情

hermes-agent

by NousResearch

10

The agent that grows with you

234,43747,175Python
AI 智能体ai-agentsagent-orchestration
查看详情

n8n

by n8n-io

12

Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.

201,88160,308TypeScript
MCP 服务器apisai-tools
查看详情

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

185,94028,768JavaScript
AI 智能体ai-agentsanthropicclaude-code
查看详情

cc-switch

by farion1231

3

A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io

128,8688,826Rust
AI 智能体claude-codeai-tools
查看详情
awesome-ai-leaderboard — Claude Code AI Skill | SkillTip