bigcode-evaluation-harness
A framework for the evaluation of autoregressive code generation language models.
👨💻 An awesome and curated list of best code-LLM for research.
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
A framework for the evaluation of autoregressive code generation language models.
A framework for the evaluation of autoregressive code generation language models on HumanEval.
A secure sandbox for running and judging code generated by LLMs.
Evaluating LLMs for Efficient Code Generation.
Measuring the LLM’s coding ability, and whether it can write new code that integrates into existing code.
BigCodeBench evaluates LLMs with practical and challenging programming tasks.
Holistic and Contamination Free Evaluation of Large Language Models for Code.
Compare performance of base multilingual code generation models on HumanEval benchmark and MultiPL-E.
BIRD contains over 12,751 unique question-SQL pairs, 95 big databases with a total size of 33.4 GB. It also covers more than 37 professional domains, such as blockchain, hockey, healthcare and education, etc.
CanAiCode Leaderboard
Coding LLMs Leaderboard
CRUXEval is a benchmark complementary to HumanEval and MBPP measuring code reasoning, understanding, and execution capabilities!
EvalPlus evaluates AI Coders with rigorous tests.
InfiBench is a comprehensive benchmark for code large language models evaluating model ability on answering freeform real-world questions in the code domain.
InterCode is a benchmark for evaluating language models on the interactive coding task. Given a natural language request, an agent is asked to interact with a software system (e.g., database, terminal) with code to resolve the issue.
They created this leaderboard to help researchers easily identify the best open-source model with an intuitive leadership quadrant graph. They evaluate the performance of open-source code models to rank them based on their capabilities and market adoption.
Spider is a large-scale complex and cross-domain semantic parsing and text-to-SQL dataset annotated by 11 Yale students. The goal of the Spider challenge is to develop natural language interfaces to cross-domain databases.
Preprint
Preprint
VoltAgent/awesome-openclaw-skills
The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞
awesome-dsh-plugin/awesome-dsh-plugin
A curated list of plugins for DeepSeek Harness (dsh) · DeepSeek Harness 插件精选列表
Kristories/awesome-guidelines
Programming style, best practices, and coding conventions.
sindresorhus/awesome
😎 Awesome lists about all kinds of interesting topics [NOTE: Pull requests are temporarily disabled until I have a chance to catch up with the existing ones]
ai-boost/awesome-prompts
Curated list of chatgpt prompts from the top-rated GPTs in the GPTs Store. Prompt Engineering, prompt attack & prompt protect. Advanced Prompt Engineering papers.
matiassingers/awesome-readme
A curated list of awesome READMEs