Skip to content
51

Awesome AI Papers

A curated list of the most impressive AI papers

1.3k stars118 forks176 entriesLast push Jul 1, 2025 (1 year ago)License none

This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.

2023 Papers >Computer Vision

Muse: Text-To-Image Generation via Masked Generative Transformers (Muse)

Structure and Content-Guided Video Synthesis with Diffusion Models (Gen-1)

ICCV, 2023

In 2 lists

Scaling Vision Transformers to 22 Billion Parameters (ViT 22B)

Adding Conditional Control to Text-to-Image Diffusion Models (ControlNet)

Lvmin Zhang

In 2 lists

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models (Visual ChatGPT)

[code]

In 3 lists

Scaling up GANs for Text-to-Image Synthesis (GigaGAN)

[Page]

In 2 lists

Segment Anything (SAM)

Alexander Kirillov

In 2 lists

DINOv2: Learning Robust Visual Features without Supervision (DINOv2)

Visual Instruction Tuning

Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models (VideoLDM)

CVPR 2023

In 2 lists

Synthetic Data from Diffusion Models Improves ImageNet Classification

Segment Anything in Medical Images (MedSAM)

Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold (DragGAN)

Neuralangelo: High-Fidelity Neural Surface Reconstruction (Neuralangelo)

SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis (SDXL)

3D Gaussian Splatting for Real-Time Radiance Field Rendering

Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization... (Qwen-VL)

MVDream: Multi-view Diffusion for 3D Generation (MVDream)

Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks (Florence-2)

VideoPoet: A Large Language Model for Zero-Shot Video Generation (VideoPoet)

2023 Papers >NLP

DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature (DetectGPT)

Toolformer: Language Models Can Teach Themselves to Use Tools (Toolformer)

by Meta AI, 2023 - A smaller model trained to translate human intention into actions (i.e. decide which APIs to call, when to call them, what arguments to pass, and how to best incorporate the results into future token prediction).

In 3 lists

LLaMA: Open and Efficient Foundation Language Models (LLaMA)

GPT-4

Announcement of GPT-4, a large multimodal model. OpenAI blog, March 14, 2023.

In 5 listsDetails

Sparks of Artificial General Intelligence: Early experiments with GPT-4 (GPT-4 Eval)

by Microsoft Research, 2023 - There are completely mind-blowing examples in the paper.

In 2 lists

HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in HuggingFace (HuggingGPT)

Solving AI Tasks with ChatGPT and its Friends in HuggingFace

In 5 listsDetails

BloombergGPT: A Large Language Model for Finance (BloombergGPT)

Instruction Tuning with GPT-4

Generative Agents: Interactive Simulacra of Human (Gen Agents)

a paper that presents computational software agents that simulate believable human behavior

In 5 listsDetails

PaLM 2 Technical Report (PaLM-2)

Tree of Thoughts: Deliberate Problem Solving with Large Language Models (ToT)

search over reasoning trees.

In 4 listsDetails

LIMA: Less Is More for Alignment (LIMA)

"less is more for alignment"; small high-quality SFT data goes a long way.

In 2 lists

QLoRA: Efficient Finetuning of Quantized LLMs (QLoRA)

Voyager: An Open-Ended Embodied Agent with Large Language Models (Voyager)

Open-ended embodied agent in Minecraft

In 3 lists

ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs (ToolLLM)

MetaGPT: Meta Programming for Multi-Agent Collaborative Framework (MetaGPT)

(2024) ICLR

In 3 lists

Code Llama: Open Foundation Models for Code (Code Llama)

Preprint

In 2 lists

RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback (RLAIF)

Large Language Models as Optimizers (OPRO)

Eureka: Human-Level Reward Design via Coding Large Language Models (Eureka)

Mathematical discoveries from program search with large language models (FunSearch)

Nature, 2024. [All Versions]. Large language models (LLMs) have demonstrated tremendous capabilities in solving complex tasks, from quantitative reasoning to understanding natural language. However, LLMs sometimes suffer from confabulations (or hallucinations), which can result in them making…

In 2 lists

2023 Papers >Audio Processing

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers (VALL-E)

MusicLM: Generating Music From Text (MusicLM)

AudioLDM: Text-to-Audio Generation with Latent Diffusion Models (AudioLDM)

Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages (USM)

Scaling Speech Technology to 1,000+ Languages (MMS)

Simple and Controllable Music Generation (MusicGen)

AudioPaLM: A Large Language Model That Can Speak and Listen (AudioPaLM)

Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale (Voicebox)

2023 Papers >Multimodal Learning

Language Is Not All You Need: Aligning Perception with Language Models (Kosmos-1)

PaLM-E: An Embodied Multimodal Language Model (PaLM-E)

AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head (AudioGPT)

Understanding and Generating Speech, Music, Sound, and Talking Head [code] [demo]

In 2 lists

ImageBind: One Embedding Space To Bind Them All (ImageBind)

CVPR'23, 2023. [All Versions]. [Project]. This work presents ImageBind, an approach to learn a joint embedding across six different modalities - images, text, audio, depth, thermal, and IMU data. The authors show that all combinations of paired data are not necessary to train such a joint…

In 2 lists

Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning (CM3Leon)

Meta-Transformer: A Unified Framework for Multimodal Learning (Meta-Transformer)

SeamlessM4T: Massively Multilingual & Multimodal Machine Translation (SeamlessM4T)

2023 Papers >Reinforcement Learning

Mastering Diverse Domains through World Models (DreamerV3)

Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, Timothy Lillicrap. Arxiv 2023; Key: DreamerV3, scaling property to world model; ExpEnv: deepmind control suite, atari, DMLab, minecraft

In 3 lists

Grounding Large Language Models in Interactive Environments with Online RL (GLAM)

Efficient Online Reinforcement Learning with Offline Data (RLPD)

Philip J. Ball, Laura Smith, Ilya Kostrikov, and Sergey Levine. ICML, 2023.

In 2 lists

Reward Design with Language Models

Direct Preference Optimization: Your Language Model is Secretly a Reward Model (DPO)

Reframed preference alignment as a simple classification objective without explicit reward modelling.

In 3 lists

Faster sorting algorithms discovered using deep reinforcement learning (AlphaDev)

Retroformer: Retrospective Large Language Agents with Policy Gradient Optimization (Retroformer)

2023 Papers >Other Papers

Symbolic Discovery of Optimization Algorithms (Lion)

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control (RT-2)

Scaling deep learning for materials discovery (GNoME)

Discovery of a structural class of antibiotics with explainable deep learning

2022 Papers >Computer Vision

A ConvNet for the 2020s (ConvNeXt)

Patches Are All You Need (ConvMixer)

(Openreview'2021)

In 2 lists

Block-NeRF: Scalable Large Scene Neural View Synthesis (Block-NeRF)

DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection (DINO)

Scaling Up Your Kernels to 31×31: Revisiting Large Kernel Design in CNNs (Large Kernel CNN)

TensoRF: Tensorial Radiance Fields (TensoRF)

MaxViT: Multi-Axis Vision Transformer (MaxViT)

Hierarchical Text-Conditional Image Generation with CLIP Latents (DALL-E 2)

Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding (Imagen)

GIT: A Generative Image-to-text Transformer for Vision and Language (GIT)

CMT: Convolutional Neural Network Meet Vision Transformers (CMT)

Swin UNETR: Swin Transformers for Semantic Segmentation of Brain Tumors... (Swin UNETR)

Classifier-Free Diffusion Guidance

Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation (DreamBooth)

DreamFusion: Text-to-3D using 2D Diffusion (DreamFusion)

. 🌐 Project Page | 💻 Code

In 2 lists

Make-A-Video: Text-to-Video Generation without Text-Video Data (Make-A-Video)

ICLR 2023

In 2 lists

On Distillation of Guided Diffusion Models

LAION-5B: An open large-scale dataset for training next generation image-text models (LAION-5B)

Imagic: Text-Based Real Image Editing with Diffusion Models (Imagic)

Visual Prompt Tuning

Magic3D: High-Resolution Text-to-3D Content Creation (Magic3D)

DiffusionDet: Diffusion Model for Object Detection (DiffusionDet)

InstructPix2Pix: Learning to Follow Image Editing Instructions (InstructPix2Pix)

Multi-Concept Customization of Text-to-Image Diffusion (Custom Diffusion)

Scalable Diffusion Models with Transformers (DiT)

2022 Papers >NLP

LaMBDA: Language Models for Dialog Applications (LaMBDA)

by Google.

In 2 lists

Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (CoT)

foundational result; intermediate reasoning steps improve performance.

In 5 listsDetails

Competition-Level Code Generation with AlphaCode (AlphaCode)

Finetuned Language Models Are Zero-Shot Learners (FLAN)

finetuned language models as zero-shot learners.

In 2 lists

Training language models to follow human instructions with human feedback (InstructGPT)

This paper presents an RLHF approach to using supervised learning to fine-tuning. It is also known as a paper that illustrates the kernel of ChatGPT's thinking. Presumably, ChatGPT is an extended version of InstructGPT that enables fine-tuning on larger datasets.

In 6 listsDetails

Multitask Prompted Training Enables Zero-Shot Task Generalization (T0)

In 2 lists

Training Compute-Optimal Large Language Models (Chinchilla)

by Hoffmann et al. at DeepMind. TLDR: introduces a new 70B LM called "Chinchilla" that outperforms much bigger LMs (GPT-3, Gopher). DeepMind has found the secret to cheaply scale large language models — to be compute-optimal, model size and training data must be scaled equally. It shows that most…

In 3 lists

Do As I Can, Not As I Say: Grounding Language in Robotic Affordances (SayCan)

GPT-NeoX-20B: An Open-Source Autoregressive Language Model (GPT-NeoX)

PaLM: Scaling Language Modeling with Pathways (PaLM)

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of lang... (BIG-bench)

Solving Quantitative Reasoning Problems with Language Models (Minerva)

[blog]

In 2 lists

ReAct: Synergizing Reasoning and Acting in Language Models (ReAct)

The foundational paper defining the Thought/Action/Observation loop structure that underlies virtually every agent harness. Required reading for understanding why the loop is structured the way it is and where each harness component maps onto the reasoning-acting cycle.

In 7 listsDetails

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model (BLOOM)

176B-parameter open multilingual LM, 46 natural languages.

In 2 lists

Optimizing Language Models for Dialogue (ChatGPT)

Announcement of ChatGPT, a conversational model trained to answer follow-up questions, admit mistakes, challenge incorrect premises, and reject inappropriate requests. OpenAI blog, November 30, 2022.

In 8 listsDetails

Large Language Models Encode Clinical Knowledge (Med-PaLM)

2022 Papers >Audio Processing

mSLAM: Massively multilingual joint pre-training for speech and text (mSLAM)

ADD 2022: the First Audio Deep Synthesis Detection Challenge (ADD)

Efficient Training of Audio Transformers with Patchout (PaSST)

MAESTRO: Matched Speech Text Representations through Modality Matching (Maestro)

SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language... (SpeechT5)

WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing (WavLM)

BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for ASR (BigSSL)

MuLan: A Joint Embedding of Music Audio and Natural Language (MuLan)

AudioLM: a Language Modeling Approach to Audio Generation (AudioLM)

AudioGen: Textually Guided Audio Generation (AudioGen)

High Fidelity Neural Audio Compression (EnCodec)

Robust Speech Recognition via Large-Scale Weak Supervision (Whisper)

2022 Papers >Multimodal Learning

BLIP: Boostrapping Language-Image Pre-training for Unified Vision-Language... (BLIP)

data2vec: A General Framework for Self-supervised Learning in Speech, Vision and... (Data2vec)

VL-Adapter: Parameter-Efficient Transfer Learning for Vision-and-Language Tasks (VL-Adapter)

Winoground: Probing Vision and Language Models for Visio-Linguistic... (Winoground)

Flamingo: a Visual Language Model for Few-Shot Learning (Flamingo)

A Generalist Agent (Gato)

CoCa: Contrastive Captioners are Image-Text Foundation Models (CoCa)

VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts (VLMo)

Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks (BEiT)

PaLI: A Jointly-Scaled Multilingual Language-Image Model (PaLI)

2022 Papers >Reinforcement Learning

Learning robust perceptive locomotion for quadrupedal robots in the wild

BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning

Outracing champion Gran Turismo drivers with deep reinforcement learning (Sophy)

Magnetic control of tokamak plasmas through deep reinforcement learning

Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning (ANYmal)

Discovering faster matrix multiplication algorithms with reinforcement learning (AlphaTensor)

2022 Papers >Other Papers

FourCastNet: A Global Data-driven High-resolution Weather Model... (FourCastNet)

ColabFold: making protein folding accessible to all (ColabFold)

Measuring and Improving the Use of Graph Information in GNN

TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis (TimesNet)

RT-1: Robotics Transformer for Real-World Control at Scale (RT-1)

Historical Papers

Perceptron: A probabilistic model for information storage and organization in the brain (Perceptron)

Learning representations by back-propagating errors (Backpropagation)

Induction of decision trees (CART)

A Tutorial on Hidden Markov Models and Selected Applications in Speech Recognition (HMM)

Multilayer feedforward networks are universal approximators

A training algorithm for optimal margin classifiers (SVM)

Bagging predictors

Gradient-based learning applied to document recognition (CNN/GTN)

Random Forests

A fast and elitist multiobjective genetic algorithm (NSGA-II)

Latent Dirichlet Allocation (LDA)

Reducing the Dimensionality of Data with Neural Networks (Autoencoder)

Visualizing Data using t-SNE (t-SNE)

In 2 lists

ImageNet: A large-scale hierarchical image database (ImageNet)

ImageNet Classification with Deep Convolutional Neural Networks (AlexNet)

Efficient Estimation of Word Representations in Vector Space (Word2vec)

Auto-Encoding Variational Bayes (VAE)

Generative Adversarial Networks (GAN)

Dropout: A Simple Way to Prevent Neural Networks from Overfitting (Dropout)

Sequence to Sequence Learning with Neural Networks

Neural Machine Translation by Jointly Learning to Align and Translate (RNNSearch-50)

This paper introduces an attention mechanism in RNNs to improve the long sequence modelling of RNNs. This paper introduces an attention mechanism to RNNs to improve their long sequence modelling capabilities. This enables RNNs to translate longer sentences more accurately.

In 2 lists

Adam: A Method for Stochastic Optimization (Adam)

Batch Normalization: Accelerating Deep Network Training by Reducing Internal Cov... (BatchNorm)

Going Deeper With Convolutions (Inception)

Human-level control through deep reinforcement learning (Deep Q Network)

Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks (Faster R-CNN)

U-Net: Convolutional Networks for Biomedical Image Segmentation (U-Net)

Deep Residual Learning for Image Recognition (ResNet)

In 2 lists

You Only Look Once: Unified, Real-Time Object Detection (YOLO)

Attention is All you Need (Transformer)

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (BERT)

bidirectional transformer pretraining; foundation for most encoder-based NLP work since 2018. Read online with section navigation and the ACL source attached.

In 5 listsDetails

Language Models are Few-Shot Learners (GPT-3)

Denoising Diffusion Probabilistic Models (DDPM)

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (ViT)

(ICLR'2021)

In 3 lists

Highly accurate protein structure prediction with AlphaFold (Alphafold)

Nature, 2021. [All Versions]. This paper provides the first computational method that can regularly predict protein structures with atomic accuracy even in cases in which no similar structure is known. This approach is a canonical application of observation- and explanation- based method for…

In 3 lists

Optimizing Language Models for Dialogue (ChatGPT)

Announcement of ChatGPT, a conversational model trained to answer follow-up questions, admit mistakes, challenge incorrect premises, and reject inappropriate requests. OpenAI blog, November 30, 2022.

In 8 listsDetails
See category
94

Table of Contents

hesreallyhim/awesome-claude-code

A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team…

Fresh★ 55k202 entriesPushed today
94

Awesome Agent Skills

VoltAgent/awesome-agent-skills

A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.

Fresh★ 35k839 entriesPushed today
93

Awesome Machine Learning

josephmisiti/awesome-machine-learning

A curated list of awesome Machine Learning frameworks, libraries and software.

Fresh★ 74k1188 entriesPushed 7 days ago
92

Awesome Production Machine Learning

EthicalML/awesome-production-machine-learning

A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning

Fresh★ 21k519 entriesPushed 3 days ago
92

AWESOME DATA SCIENCE

academic/awesome-datascience

:memo: An awesome Data Science repository to learn and apply for real world problems.

Fresh★ 30k881 entriesPushed today
91

Static Analysis

analysis-tools-dev/static-analysis

⚙️ A curated list of static analysis (SAST) tools and linters for all programming languages, config files, build tools, and more. The focus is on tools which improve…

Fresh★ 15k528 entriesPushed 8 days ago